* [PATCH v7 01/24] drm/xe: reference VM from PT BOs
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
@ 2026-09-25 4:52 ` Matthew Brost
2026-09-25 12:00 ` Francois Dugast
2026-09-25 4:52 ` [PATCH v7 02/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
` (27 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:52 UTC (permalink / raw)
To: intel-xe
PT BOs share the VM's dma-resv and therefore must keep the VM alive
until all PT BOs are destroyed.
Take a VM reference from PT BOs to ensure the shared dma-resv remains
valid for the lifetime of the BOs. Scope the change to PT BOs rather
than all kernel BOs to minimize risk and avoid extending VM lifetimes
unnecessarily.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
---
drivers/gpu/drm/xe/xe_bo.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
index f2ab9bf43a86..6921b6967330 100644
--- a/drivers/gpu/drm/xe/xe_bo.c
+++ b/drivers/gpu/drm/xe/xe_bo.c
@@ -1879,7 +1879,7 @@ static void xe_ttm_bo_destroy(struct ttm_buffer_object *ttm_bo)
xe_drm_client_remove_bo(bo);
#endif
- if (bo->vm && xe_bo_is_user(bo))
+ if (bo->vm && (xe_bo_is_user(bo) || bo->flags & XE_BO_FLAG_PAGETABLE))
xe_vm_put(bo->vm);
if (bo->parent_obj)
@@ -2575,7 +2575,7 @@ __xe_bo_create_locked(struct xe_device *xe,
* by having all the vm's bo refereferences released at vm close
* time.
*/
- if (vm && xe_bo_is_user(bo))
+ if (vm && (xe_bo_is_user(bo) || bo->flags & XE_BO_FLAG_PAGETABLE))
xe_vm_get(vm);
bo->vm = vm;
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 01/24] drm/xe: reference VM from PT BOs
2026-09-25 4:52 ` [PATCH v7 01/24] drm/xe: reference VM from PT BOs Matthew Brost
@ 2026-09-25 12:00 ` Francois Dugast
2026-09-25 16:10 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: Francois Dugast @ 2026-09-25 12:00 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Nit: first letter of the commit title is usually capitalized.
On Thu, Sep 24, 2026 at 09:52:57PM -0700, Matthew Brost wrote:
> PT BOs share the VM's dma-resv and therefore must keep the VM alive
> until all PT BOs are destroyed.
>
> Take a VM reference from PT BOs to ensure the shared dma-resv remains
> valid for the lifetime of the BOs. Scope the change to PT BOs rather
> than all kernel BOs to minimize risk and avoid extending VM lifetimes
> unnecessarily.
>
> Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Did you observe any side effect without this change in earlier versions
of this series?
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
> ---
> drivers/gpu/drm/xe/xe_bo.c | 4 ++--
> 1 file changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
> index f2ab9bf43a86..6921b6967330 100644
> --- a/drivers/gpu/drm/xe/xe_bo.c
> +++ b/drivers/gpu/drm/xe/xe_bo.c
> @@ -1879,7 +1879,7 @@ static void xe_ttm_bo_destroy(struct ttm_buffer_object *ttm_bo)
> xe_drm_client_remove_bo(bo);
> #endif
>
> - if (bo->vm && xe_bo_is_user(bo))
> + if (bo->vm && (xe_bo_is_user(bo) || bo->flags & XE_BO_FLAG_PAGETABLE))
> xe_vm_put(bo->vm);
>
> if (bo->parent_obj)
> @@ -2575,7 +2575,7 @@ __xe_bo_create_locked(struct xe_device *xe,
> * by having all the vm's bo refereferences released at vm close
> * time.
> */
> - if (vm && xe_bo_is_user(bo))
> + if (vm && (xe_bo_is_user(bo) || bo->flags & XE_BO_FLAG_PAGETABLE))
> xe_vm_get(vm);
> bo->vm = vm;
>
> --
> 2.34.1
>
^ permalink raw reply [flat|nested] 54+ messages in thread
* Re: [PATCH v7 01/24] drm/xe: reference VM from PT BOs
2026-09-25 12:00 ` Francois Dugast
@ 2026-09-25 16:10 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 16:10 UTC (permalink / raw)
To: Francois Dugast; +Cc: intel-xe
On Fri, Sep 25, 2026 at 02:00:22PM +0200, Francois Dugast wrote:
> Nit: first letter of the commit title is usually capitalized.
>
> On Thu, Sep 24, 2026 at 09:52:57PM -0700, Matthew Brost wrote:
> > PT BOs share the VM's dma-resv and therefore must keep the VM alive
> > until all PT BOs are destroyed.
> >
> > Take a VM reference from PT BOs to ensure the shared dma-resv remains
> > valid for the lifetime of the BOs. Scope the change to PT BOs rather
> > than all kernel BOs to minimize risk and avoid extending VM lifetimes
> > unnecessarily.
> >
> > Signed-off-by: Matthew Brost <matthew.brost@intel.com>
>
> Did you observe any side effect without this change in earlier versions
> of this series?
No, just Sashiko (correctly) complaining. The kmalloc in
individualization step would have to fail to get a UAF, thus pretty hard
for issues to show up. We might have issues here in other existing code
path if that happens but wanted to keep this change minimal rather than
auditing the entire driver.
Again Christian's series to ref count dma-resv is the real fix for
memory safety around dma-resv sharing. Hopefully we gets that in soon -
I basically gave him my ack to merge it but subsystem wide change so
might get held up.
Matt
>
> Reviewed-by: Francois Dugast <francois.dugast@intel.com>
>
> > ---
> > drivers/gpu/drm/xe/xe_bo.c | 4 ++--
> > 1 file changed, 2 insertions(+), 2 deletions(-)
> >
> > diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
> > index f2ab9bf43a86..6921b6967330 100644
> > --- a/drivers/gpu/drm/xe/xe_bo.c
> > +++ b/drivers/gpu/drm/xe/xe_bo.c
> > @@ -1879,7 +1879,7 @@ static void xe_ttm_bo_destroy(struct ttm_buffer_object *ttm_bo)
> > xe_drm_client_remove_bo(bo);
> > #endif
> >
> > - if (bo->vm && xe_bo_is_user(bo))
> > + if (bo->vm && (xe_bo_is_user(bo) || bo->flags & XE_BO_FLAG_PAGETABLE))
> > xe_vm_put(bo->vm);
> >
> > if (bo->parent_obj)
> > @@ -2575,7 +2575,7 @@ __xe_bo_create_locked(struct xe_device *xe,
> > * by having all the vm's bo refereferences released at vm close
> > * time.
> > */
> > - if (vm && xe_bo_is_user(bo))
> > + if (vm && (xe_bo_is_user(bo) || bo->flags & XE_BO_FLAG_PAGETABLE))
> > xe_vm_get(vm);
> > bo->vm = vm;
> >
> > --
> > 2.34.1
> >
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 02/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
2026-09-25 4:52 ` [PATCH v7 01/24] drm/xe: reference VM from PT BOs Matthew Brost
@ 2026-09-25 4:52 ` Matthew Brost
2026-09-25 4:52 ` [PATCH v7 03/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
` (26 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:52 UTC (permalink / raw)
To: intel-xe; +Cc: Francois Dugast
Remove the xe_migrate_pt_update argument from the populate and clear
vfuns. This structure will not be available in run_job, where CPU binds
will be implemented. The populate path no longer needs it, and the clear
path already uses the VM field instead.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
---
drivers/gpu/drm/xe/xe_migrate.c | 9 +++++----
drivers/gpu/drm/xe/xe_migrate.h | 12 +++++-------
drivers/gpu/drm/xe/xe_pt.c | 12 +++++-------
3 files changed, 15 insertions(+), 18 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index ff45c24d8889..149c5fa654e6 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -1760,6 +1760,7 @@ static void write_pgtable(struct xe_tile *tile, struct xe_bb *bb, u64 ppgtt_ofs,
struct xe_migrate_pt_update *pt_update)
{
const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
+ struct xe_vm *vm = pt_update->vops->vm;
u32 chunk;
u32 ofs = update->ofs, size = update->qwords;
@@ -1791,10 +1792,10 @@ static void write_pgtable(struct xe_tile *tile, struct xe_bb *bb, u64 ppgtt_ofs,
bb->cs[bb->len++] = lower_32_bits(addr);
bb->cs[bb->len++] = upper_32_bits(addr);
if (pt_op->bind)
- ops->populate(pt_update, tile, NULL, bb->cs + bb->len,
+ ops->populate(tile, NULL, bb->cs + bb->len,
ofs, chunk, update);
else
- ops->clear(pt_update, tile, NULL, bb->cs + bb->len,
+ ops->clear(vm, tile, NULL, bb->cs + bb->len,
ofs, chunk, update);
bb->len += chunk * 2;
@@ -1851,12 +1852,12 @@ xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
&pt_op->entries[j];
if (pt_op->bind)
- ops->populate(pt_update, m->tile,
+ ops->populate(m->tile,
&update->pt_bo->vmap, NULL,
update->ofs, update->qwords,
update);
else
- ops->clear(pt_update, m->tile,
+ ops->clear(vm, m->tile,
&update->pt_bo->vmap, NULL,
update->ofs, update->qwords, update);
}
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index a9acc62f78f0..2ec9de896dfe 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -40,7 +40,6 @@ enum xe_migrate_copy_dir {
struct xe_migrate_pt_update_ops {
/**
* @populate: Populate a command buffer or page-table with ptes.
- * @pt_update: Embeddable callback argument.
* @tile: The tile for the current operation.
* @map: struct iosys_map into the memory to be populated.
* @pos: If @map is NULL, map into the memory to be populated.
@@ -52,13 +51,12 @@ struct xe_migrate_pt_update_ops {
* page-table system to populate command buffers or shared
* page-tables with PTEs.
*/
- void (*populate)(struct xe_migrate_pt_update *pt_update,
- struct xe_tile *tile, struct iosys_map *map,
+ void (*populate)(struct xe_tile *tile, struct iosys_map *map,
void *pos, u32 ofs, u32 num_qwords,
const struct xe_vm_pgtable_update *update);
/**
* @clear: Clear a command buffer or page-table with ptes.
- * @pt_update: Embeddable callback argument.
+ * @vm: VM being updated
* @tile: The tile for the current operation.
* @map: struct iosys_map into the memory to be populated.
* @pos: If @map is NULL, map into the memory to be populated.
@@ -70,9 +68,9 @@ struct xe_migrate_pt_update_ops {
* page-table system to populate command buffers or shared
* page-tables with PTEs.
*/
- void (*clear)(struct xe_migrate_pt_update *pt_update,
- struct xe_tile *tile, struct iosys_map *map,
- void *pos, u32 ofs, u32 num_qwords,
+ void (*clear)(struct xe_vm *vm, struct xe_tile *tile,
+ struct iosys_map *map, void *pos, u32 ofs,
+ u32 num_qwords,
const struct xe_vm_pgtable_update *update);
/**
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index 4b351dbf6572..dd007e3bc40e 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -1102,9 +1102,8 @@ bool xe_pt_zap_ptes_range(struct xe_tile *tile, struct xe_vm *vm,
}
static void
-xe_vm_populate_pgtable(struct xe_migrate_pt_update *pt_update, struct xe_tile *tile,
- struct iosys_map *map, void *data,
- u32 qword_ofs, u32 num_qwords,
+xe_vm_populate_pgtable(struct xe_tile *tile, struct iosys_map *map,
+ void *data, u32 qword_ofs, u32 num_qwords,
const struct xe_vm_pgtable_update *update)
{
struct xe_pt_entry *ptes = update->pt_entries;
@@ -2018,12 +2017,11 @@ static unsigned int xe_pt_stage_unbind(struct xe_tile *tile,
}
static void
-xe_migrate_clear_pgtable_callback(struct xe_migrate_pt_update *pt_update,
- struct xe_tile *tile, struct iosys_map *map,
- void *ptr, u32 qword_ofs, u32 num_qwords,
+xe_migrate_clear_pgtable_callback(struct xe_vm *vm, struct xe_tile *tile,
+ struct iosys_map *map, void *ptr,
+ u32 qword_ofs, u32 num_qwords,
const struct xe_vm_pgtable_update *update)
{
- struct xe_vm *vm = pt_update->vops->vm;
u64 empty = __xe_pt_empty_pte(tile, vm, update->pt->level);
int i;
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 03/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
2026-09-25 4:52 ` [PATCH v7 01/24] drm/xe: reference VM from PT BOs Matthew Brost
2026-09-25 4:52 ` [PATCH v7 02/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
@ 2026-09-25 4:52 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 04/24] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
` (25 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:52 UTC (permalink / raw)
To: intel-xe; +Cc: Francois Dugast
Add the xe_migrate_update_pgtables_cpu_execute helper, which performs
the CPU-side page-table update. This will support implementing CPU
binds, as the submission backend can call this helper once a bind job’s
dependencies are resolved. While here, add assertions to provide basic
sanity checks on tht function arguments.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
---
v7:
- Fix assert structure (Sashiko)
---
drivers/gpu/drm/xe/xe_migrate.c | 56 +++++++++++++++++++--------------
1 file changed, 33 insertions(+), 23 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 149c5fa654e6..40e60567cb58 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -1819,6 +1819,36 @@ struct migrate_test_params {
container_of(_priv, struct migrate_test_params, base)
#endif
+static void
+xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
+ const struct xe_migrate_pt_update_ops *ops,
+ struct xe_vm_pgtable_update_op *pt_op,
+ u32 num_ops)
+{
+ u32 j, i;
+
+ for (j = 0; j < num_ops; ++j, ++pt_op) {
+ for (i = 0; i < pt_op->num_entries; i++) {
+ const struct xe_vm_pgtable_update *update =
+ &pt_op->entries[i];
+
+ xe_tile_assert(tile, !iosys_map_is_null(&update->pt_bo->vmap));
+
+ if (pt_op->bind)
+ ops->populate(tile, &update->pt_bo->vmap,
+ NULL, update->ofs, update->qwords,
+ update);
+ else
+ ops->clear(vm, tile, &update->pt_bo->vmap,
+ NULL, update->ofs, update->qwords,
+ update);
+ }
+ }
+
+ trace_xe_vm_cpu_bind(vm);
+ xe_device_wmb(vm->xe);
+}
+
static struct dma_fence *
xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
struct xe_migrate_pt_update *pt_update)
@@ -1831,7 +1861,6 @@ xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
struct xe_vm_pgtable_update_ops *pt_update_ops =
&pt_update->vops->pt_update_ops[pt_update->tile_id];
int err;
- u32 i, j;
if (XE_TEST_ONLY(test && test->force_gpu))
return ERR_PTR(-ETIME);
@@ -1843,28 +1872,9 @@ xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
return ERR_PTR(err);
}
- for (i = 0; i < pt_update_ops->num_ops; ++i) {
- const struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->ops[i];
-
- for (j = 0; j < pt_op->num_entries; j++) {
- const struct xe_vm_pgtable_update *update =
- &pt_op->entries[j];
-
- if (pt_op->bind)
- ops->populate(m->tile,
- &update->pt_bo->vmap, NULL,
- update->ofs, update->qwords,
- update);
- else
- ops->clear(vm, m->tile,
- &update->pt_bo->vmap, NULL,
- update->ofs, update->qwords, update);
- }
- }
-
- trace_xe_vm_cpu_bind(vm);
- xe_device_wmb(vm->xe);
+ xe_migrate_update_pgtables_cpu_execute(vm, m->tile, ops,
+ pt_update_ops->ops,
+ pt_update_ops->num_ops);
return dma_fence_get_stub();
}
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 04/24] drm/xe: Decouple exec queue idle check from LRC
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (2 preceding siblings ...)
2026-09-25 4:52 ` [PATCH v7 03/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 05/24] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
` (24 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Stuart Summers
We already maintain a job count for each exec queue, so simplify the idle
check to rely on the job count rather than the LRC state. This decouples
exec queues from LRC-based backends and avoids unnecessarily coupling idle
detection to backend-specific implementation details.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Stuart Summers <stuart.summers@intel.com>
---
drivers/gpu/drm/xe/xe_exec_queue.c | 15 +--------------
1 file changed, 1 insertion(+), 14 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.c b/drivers/gpu/drm/xe/xe_exec_queue.c
index e63559a2f582..7bbb31431900 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.c
+++ b/drivers/gpu/drm/xe/xe_exec_queue.c
@@ -1580,20 +1580,7 @@ bool xe_exec_queue_is_lr(struct xe_exec_queue *q)
*/
bool xe_exec_queue_is_idle(struct xe_exec_queue *q)
{
- if (xe_exec_queue_is_parallel(q)) {
- int i;
-
- for (i = 0; i < q->width; ++i) {
- if (xe_lrc_seqno(q->lrc[i]) !=
- q->lrc[i]->fence_ctx.next_seqno - 1)
- return false;
- }
-
- return true;
- }
-
- return xe_lrc_seqno(q->lrc[0]) ==
- q->lrc[0]->fence_ctx.next_seqno - 1;
+ return !atomic_read(&q->job_cnt);
}
/**
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 05/24] drm/xe: Add job count to GuC exec queue snapshot
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (3 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 04/24] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 06/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
` (23 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Stuart Summers
Add the job count to the GuC exec queue snapshot, as this is useful
debug information.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Stuart Summers <stuart.summers@intel.com>
---
v7:
- s/%d/%u (Sashiko)
---
drivers/gpu/drm/xe/xe_guc_submit.c | 2 ++
drivers/gpu/drm/xe/xe_guc_submit_types.h | 2 ++
2 files changed, 4 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index f3ba8abfc228..4bd1ead57efa 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -3686,6 +3686,7 @@ xe_guc_exec_queue_snapshot_capture(struct xe_exec_queue *q)
snapshot->logical_mask = q->logical_mask;
snapshot->width = q->width;
snapshot->refcount = kref_read(&q->refcount);
+ snapshot->jobcount = atomic_read(&q->job_cnt);
snapshot->sched_timeout = sched->base.timeout;
snapshot->sched_props.timeslice_us = q->sched_props.timeslice_us;
snapshot->sched_props.preempt_timeout_us =
@@ -3758,6 +3759,7 @@ xe_guc_exec_queue_snapshot_print(struct xe_guc_submit_exec_queue_snapshot *snaps
drm_printf(p, "\tLogical mask: 0x%x\n", snapshot->logical_mask);
drm_printf(p, "\tWidth: %d\n", snapshot->width);
drm_printf(p, "\tRef: %d\n", snapshot->refcount);
+ drm_printf(p, "\tJob count: %u\n", snapshot->jobcount);
drm_printf(p, "\tTimeout: %ld (ms)\n", snapshot->sched_timeout);
drm_printf(p, "\tTimeslice: %u (us)\n",
snapshot->sched_props.timeslice_us);
diff --git a/drivers/gpu/drm/xe/xe_guc_submit_types.h b/drivers/gpu/drm/xe/xe_guc_submit_types.h
index 7824f61b1290..8271702e692e 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit_types.h
+++ b/drivers/gpu/drm/xe/xe_guc_submit_types.h
@@ -77,6 +77,8 @@ struct xe_guc_submit_exec_queue_snapshot {
u16 width;
/** @refcount: ref count of this exec queue */
u32 refcount;
+ /** @jobcount: job count of this exec queue */
+ u32 jobcount;
/**
* @sched_timeout: the time after which a job is removed from the
* scheduler.
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 06/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (4 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 05/24] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
` (22 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Francois Dugast
Update the xe_bo_put_deferred arguments to include a writeback flag,
which indicates whether the BO was added to the deferred list. This is
useful when the caller needs to take additional actions after the BO has
been queued for deferred release.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
---
drivers/gpu/drm/xe/xe_bo.h | 10 ++++++++--
drivers/gpu/drm/xe/xe_drm_client.c | 2 +-
drivers/gpu/drm/xe/xe_pt.c | 2 +-
3 files changed, 10 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_bo.h b/drivers/gpu/drm/xe/xe_bo.h
index 341fa93a71e4..861b1be231de 100644
--- a/drivers/gpu/drm/xe/xe_bo.h
+++ b/drivers/gpu/drm/xe/xe_bo.h
@@ -496,6 +496,8 @@ void __xe_bo_release_dummy(struct kref *kref);
* @bo: The bo to put.
* @deferred: List to which to add the buffer object if we cannot put, or
* NULL if the function is to put unconditionally.
+ * @added: BO was added to deferred list, written back to caller, can be NULL if
+ * writeback is not needed. Only set to true when added, never set to false.
*
* Since the final freeing of an object includes both sleeping and (!)
* memory allocation in the dma_resv individualization, it's not ok
@@ -515,7 +517,8 @@ void __xe_bo_release_dummy(struct kref *kref);
* false otherwise.
*/
static inline bool
-xe_bo_put_deferred(struct xe_bo *bo, struct llist_head *deferred)
+xe_bo_put_deferred(struct xe_bo *bo, struct llist_head *deferred,
+ bool *added)
{
if (!deferred) {
xe_bo_put(bo);
@@ -525,6 +528,9 @@ xe_bo_put_deferred(struct xe_bo *bo, struct llist_head *deferred)
if (!kref_put(&bo->ttm.base.refcount, __xe_bo_release_dummy))
return false;
+ if (added)
+ *added = true;
+
return llist_add(&bo->freed, deferred);
}
@@ -541,7 +547,7 @@ xe_bo_put_async(struct xe_bo *bo)
{
struct xe_bo_dev *bo_device = &xe_bo_device(bo)->bo_device;
- if (xe_bo_put_deferred(bo, &bo_device->async_list))
+ if (xe_bo_put_deferred(bo, &bo_device->async_list, NULL))
schedule_work(&bo_device->async_free);
}
diff --git a/drivers/gpu/drm/xe/xe_drm_client.c b/drivers/gpu/drm/xe/xe_drm_client.c
index e116fb562c4c..4c424d1c6721 100644
--- a/drivers/gpu/drm/xe/xe_drm_client.c
+++ b/drivers/gpu/drm/xe/xe_drm_client.c
@@ -256,7 +256,7 @@ static void show_meminfo(struct drm_printer *p, struct drm_file *file)
xe_assert(xef->xe, !list_empty(&bo->client_link));
}
- xe_bo_put_deferred(bo, &deferred);
+ xe_bo_put_deferred(bo, &deferred, NULL);
}
spin_unlock(&client->bos_lock);
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index dd007e3bc40e..42a37e40a6c0 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -213,7 +213,7 @@ void xe_pt_destroy(struct xe_pt *pt, u32 flags, struct llist_head *deferred)
XE_WARN_ON(!list_empty(&pt->bo->ttm.base.gpuva.list));
xe_bo_unpin(pt->bo);
- xe_bo_put_deferred(pt->bo, deferred);
+ xe_bo_put_deferred(pt->bo, deferred, NULL);
if (pt->level > 0 && pt->num_live) {
struct xe_pt_dir *pt_dir = as_xe_pt_dir(pt);
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (5 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 06/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 11:23 ` Francois Dugast
2026-09-25 4:53 ` [PATCH v7 08/24] drm/xe: Add helpers to access PT ops Matthew Brost
` (21 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe
Update the scheduler job layer to support PT jobs. PT jobs are executed
entirely on the CPU and do not require LRC fences or a batch address.
Repurpose the LRC fence storage to hold PT‑job arguments and update the
scheduler job layer to distinguish between PT jobs and jobs that require
an LRC.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
---
v7:
- Move is_pt_job goto above seqno/tlb flush code (Sashiko, Francois)
---
drivers/gpu/drm/xe/xe_sched_job.c | 95 ++++++++++++++++---------
drivers/gpu/drm/xe/xe_sched_job_types.h | 31 +++++++-
drivers/gpu/drm/xe/xe_trace.h | 2 +-
3 files changed, 92 insertions(+), 36 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_sched_job.c b/drivers/gpu/drm/xe/xe_sched_job.c
index a4fa00632a30..20992ea816e6 100644
--- a/drivers/gpu/drm/xe/xe_sched_job.c
+++ b/drivers/gpu/drm/xe/xe_sched_job.c
@@ -26,19 +26,22 @@ static struct kmem_cache *xe_sched_job_parallel_slab;
int __init xe_sched_job_module_init(void)
{
+ struct xe_sched_job *job;
+ size_t size;
+
+ size = struct_size(job, ptrs, 1);
xe_sched_job_slab =
- kmem_cache_create("xe_sched_job",
- sizeof(struct xe_sched_job) +
- sizeof(struct xe_job_ptrs), 0,
+ kmem_cache_create("xe_sched_job", size, 0,
SLAB_HWCACHE_ALIGN, NULL);
if (!xe_sched_job_slab)
return -ENOMEM;
+ size = max_t(size_t,
+ struct_size(job, ptrs,
+ XE_HW_ENGINE_MAX_INSTANCE),
+ struct_size(job, pt_update, 1));
xe_sched_job_parallel_slab =
- kmem_cache_create("xe_sched_job_parallel",
- sizeof(struct xe_sched_job) +
- sizeof(struct xe_job_ptrs) *
- XE_HW_ENGINE_MAX_INSTANCE, 0,
+ kmem_cache_create("xe_sched_job_parallel", size, 0,
SLAB_HWCACHE_ALIGN, NULL);
if (!xe_sched_job_parallel_slab) {
kmem_cache_destroy(xe_sched_job_slab);
@@ -84,6 +87,9 @@ static void xe_sched_job_free_fences(struct xe_sched_job *job)
{
int i;
+ if (job->is_pt_job)
+ return;
+
for (i = 0; i < job->q->width; ++i) {
struct xe_job_ptrs *ptrs = &job->ptrs[i];
@@ -93,10 +99,23 @@ static void xe_sched_job_free_fences(struct xe_sched_job *job)
}
}
+/**
+ * xe_sched_job_create() - Create a scheduler job
+ * @q: exec queue to create the scheduler job for
+ * @batch_addr: array of batch addresses for the job; must match the width of
+ * @q, or NULL to indicate a PT job that does not require a batch address
+ *
+ * Create a scheduler job for submission.
+ *
+ * Context: Reclaim
+ *
+ * Return: a &xe_sched_job object on success, or an ERR_PTR on failure.
+ */
struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
u64 *batch_addr)
{
bool is_migration = xe_sched_job_is_migration(q);
+ struct xe_device *xe = gt_to_xe(q->gt);
struct xe_sched_job *job;
int err;
int i;
@@ -105,6 +124,9 @@ struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
/* only a kernel context can submit a vm-less job */
XE_WARN_ON(!q->vm && !(q->flags & EXEC_QUEUE_FLAG_KERNEL));
+ xe_assert(xe, batch_addr ||
+ q->flags & (EXEC_QUEUE_FLAG_VM | EXEC_QUEUE_FLAG_MIGRATE));
+
job = job_alloc(xe_exec_queue_is_parallel(q) || is_migration);
if (!job)
return ERR_PTR(-ENOMEM);
@@ -119,34 +141,39 @@ struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
if (err)
goto err_free;
- for (i = 0; i < q->width; ++i) {
- struct dma_fence *fence = xe_lrc_alloc_seqno_fence();
- struct dma_fence_chain *chain;
-
- if (IS_ERR(fence)) {
- err = PTR_ERR(fence);
- goto err_sched_job;
+ if (!batch_addr) {
+ job->fence = dma_fence_get_stub();
+ job->is_pt_job = true;
+ } else {
+ for (i = 0; i < q->width; ++i) {
+ struct dma_fence *fence = xe_lrc_alloc_seqno_fence();
+ struct dma_fence_chain *chain;
+
+ if (IS_ERR(fence)) {
+ err = PTR_ERR(fence);
+ goto err_sched_job;
+ }
+ job->ptrs[i].lrc_fence = fence;
+
+ if (i + 1 == q->width)
+ continue;
+
+ chain = dma_fence_chain_alloc();
+ if (!chain) {
+ err = -ENOMEM;
+ goto err_sched_job;
+ }
+ job->ptrs[i].chain_fence = chain;
}
- job->ptrs[i].lrc_fence = fence;
- if (i + 1 == q->width)
- continue;
+ width = q->width;
+ if (is_migration)
+ width = 2;
- chain = dma_fence_chain_alloc();
- if (!chain) {
- err = -ENOMEM;
- goto err_sched_job;
- }
- job->ptrs[i].chain_fence = chain;
+ for (i = 0; i < width; ++i)
+ job->ptrs[i].batch_addr = batch_addr[i];
}
- width = q->width;
- if (is_migration)
- width = 2;
-
- for (i = 0; i < width; ++i)
- job->ptrs[i].batch_addr = batch_addr[i];
-
atomic_inc(&q->job_cnt);
xe_pm_runtime_get_noresume(job_to_xe(job));
trace_xe_sched_job_create(job);
@@ -246,7 +273,7 @@ bool xe_sched_job_completed(struct xe_sched_job *job)
void xe_sched_job_arm(struct xe_sched_job *job)
{
struct xe_exec_queue *q = job->q;
- struct dma_fence *fence, *prev;
+ struct dma_fence *fence = job->fence, *prev;
struct xe_vm *vm = q->vm;
u64 seqno = 0;
int i;
@@ -259,6 +286,9 @@ void xe_sched_job_arm(struct xe_sched_job *job)
xe_vm_assert_held(q->vm);
}
+ if (job->is_pt_job)
+ goto arm;
+
if (vm && !xe_sched_job_is_migration(q) && !xe_vm_in_lr_mode(vm) &&
(vm->batch_invalidate_tlb || vm->tlb_flush_seqno != q->tlb_flush_seqno)) {
xe_vm_assert_held(vm);
@@ -286,6 +316,7 @@ void xe_sched_job_arm(struct xe_sched_job *job)
fence = &chain->base;
}
+arm:
job->fence = dma_fence_get(fence); /* Pairs with put in scheduler */
drm_sched_job_arm(&job->drm);
}
@@ -329,7 +360,7 @@ xe_sched_job_snapshot_capture(struct xe_sched_job *job)
snapshot->batch_addr_len = q->width;
for (i = 0; i < q->width; i++)
- snapshot->batch_addr[i] =
+ snapshot->batch_addr[i] = job->is_pt_job ? 0 :
xe_device_uncanonicalize_addr(xe, job->ptrs[i].batch_addr);
return snapshot;
diff --git a/drivers/gpu/drm/xe/xe_sched_job_types.h b/drivers/gpu/drm/xe/xe_sched_job_types.h
index 0490b1247a6e..5e1824c36c74 100644
--- a/drivers/gpu/drm/xe/xe_sched_job_types.h
+++ b/drivers/gpu/drm/xe/xe_sched_job_types.h
@@ -10,10 +10,29 @@
#include <drm/gpu_scheduler.h>
-struct xe_exec_queue;
struct dma_fence;
struct dma_fence_chain;
+struct xe_exec_queue;
+struct xe_migrate_pt_update_ops;
+struct xe_pt_job_ops;
+struct xe_tile;
+struct xe_vm;
+
+/**
+ * struct xe_pt_update_args - PT update arguments
+ */
+struct xe_pt_update_args {
+ /** @vm: VM which is being bound */
+ struct xe_vm *vm;
+ /** @tile: Tile which page tables belong to */
+ struct xe_tile *tile;
+ /** @ops: Migrate PT update ops */
+ const struct xe_migrate_pt_update_ops *ops;
+ /** @pt_job_ops: PT job ops state */
+ struct xe_pt_job_ops *pt_job_ops;
+};
+
/**
* struct xe_job_ptrs - Per hw engine instance data
*/
@@ -71,8 +90,14 @@ struct xe_sched_job {
bool restore_replay;
/** @last_replay: last job being replayed */
bool last_replay;
- /** @ptrs: per instance pointers. */
- struct xe_job_ptrs ptrs[];
+ /** @is_pt_job: is a PT job */
+ bool is_pt_job;
+ union {
+ /** @ptrs: per instance pointers. */
+ DECLARE_FLEX_ARRAY(struct xe_job_ptrs, ptrs);
+ /** @pt_update: PT update arguments */
+ DECLARE_FLEX_ARRAY(struct xe_pt_update_args, pt_update);
+ };
};
struct xe_sched_job_snapshot {
diff --git a/drivers/gpu/drm/xe/xe_trace.h b/drivers/gpu/drm/xe/xe_trace.h
index 2fe8f89a1e34..d4e9d91f6f7f 100644
--- a/drivers/gpu/drm/xe/xe_trace.h
+++ b/drivers/gpu/drm/xe/xe_trace.h
@@ -261,7 +261,7 @@ DECLARE_EVENT_CLASS(xe_sched_job,
__entry->flags = job->q->flags;
__entry->error = job->fence ? job->fence->error : 0;
__entry->fence = job->fence;
- __entry->batch_addr = (u64)job->ptrs[0].batch_addr;
+ __entry->batch_addr = job->is_pt_job ? 0 : (u64)job->ptrs[0].batch_addr;
),
TP_printk("dev=%s, fence=%p, seqno=%u, lrc_seqno=%u, gt=%u, guc_id=%d, batch_addr=0x%012llx, guc_state=0x%x, flags=0x%x, error=%d",
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs
2026-09-25 4:53 ` [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
@ 2026-09-25 11:23 ` Francois Dugast
0 siblings, 0 replies; 54+ messages in thread
From: Francois Dugast @ 2026-09-25 11:23 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
On Thu, Sep 24, 2026 at 09:53:03PM -0700, Matthew Brost wrote:
> Update the scheduler job layer to support PT jobs. PT jobs are executed
> entirely on the CPU and do not require LRC fences or a batch address.
> Repurpose the LRC fence storage to hold PT‑job arguments and update the
> scheduler job layer to distinguish between PT jobs and jobs that require
> an LRC.
>
> Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
>
> ---
> v7:
> - Move is_pt_job goto above seqno/tlb flush code (Sashiko, Francois)
> ---
> drivers/gpu/drm/xe/xe_sched_job.c | 95 ++++++++++++++++---------
> drivers/gpu/drm/xe/xe_sched_job_types.h | 31 +++++++-
> drivers/gpu/drm/xe/xe_trace.h | 2 +-
> 3 files changed, 92 insertions(+), 36 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_sched_job.c b/drivers/gpu/drm/xe/xe_sched_job.c
> index a4fa00632a30..20992ea816e6 100644
> --- a/drivers/gpu/drm/xe/xe_sched_job.c
> +++ b/drivers/gpu/drm/xe/xe_sched_job.c
> @@ -26,19 +26,22 @@ static struct kmem_cache *xe_sched_job_parallel_slab;
>
> int __init xe_sched_job_module_init(void)
> {
> + struct xe_sched_job *job;
> + size_t size;
> +
> + size = struct_size(job, ptrs, 1);
> xe_sched_job_slab =
> - kmem_cache_create("xe_sched_job",
> - sizeof(struct xe_sched_job) +
> - sizeof(struct xe_job_ptrs), 0,
> + kmem_cache_create("xe_sched_job", size, 0,
> SLAB_HWCACHE_ALIGN, NULL);
> if (!xe_sched_job_slab)
> return -ENOMEM;
>
> + size = max_t(size_t,
> + struct_size(job, ptrs,
> + XE_HW_ENGINE_MAX_INSTANCE),
> + struct_size(job, pt_update, 1));
> xe_sched_job_parallel_slab =
> - kmem_cache_create("xe_sched_job_parallel",
> - sizeof(struct xe_sched_job) +
> - sizeof(struct xe_job_ptrs) *
> - XE_HW_ENGINE_MAX_INSTANCE, 0,
> + kmem_cache_create("xe_sched_job_parallel", size, 0,
> SLAB_HWCACHE_ALIGN, NULL);
> if (!xe_sched_job_parallel_slab) {
> kmem_cache_destroy(xe_sched_job_slab);
> @@ -84,6 +87,9 @@ static void xe_sched_job_free_fences(struct xe_sched_job *job)
> {
> int i;
>
> + if (job->is_pt_job)
> + return;
> +
> for (i = 0; i < job->q->width; ++i) {
> struct xe_job_ptrs *ptrs = &job->ptrs[i];
>
> @@ -93,10 +99,23 @@ static void xe_sched_job_free_fences(struct xe_sched_job *job)
> }
> }
>
> +/**
> + * xe_sched_job_create() - Create a scheduler job
> + * @q: exec queue to create the scheduler job for
> + * @batch_addr: array of batch addresses for the job; must match the width of
> + * @q, or NULL to indicate a PT job that does not require a batch address
> + *
> + * Create a scheduler job for submission.
> + *
> + * Context: Reclaim
> + *
> + * Return: a &xe_sched_job object on success, or an ERR_PTR on failure.
> + */
> struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
> u64 *batch_addr)
> {
> bool is_migration = xe_sched_job_is_migration(q);
> + struct xe_device *xe = gt_to_xe(q->gt);
> struct xe_sched_job *job;
> int err;
> int i;
> @@ -105,6 +124,9 @@ struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
> /* only a kernel context can submit a vm-less job */
> XE_WARN_ON(!q->vm && !(q->flags & EXEC_QUEUE_FLAG_KERNEL));
>
> + xe_assert(xe, batch_addr ||
> + q->flags & (EXEC_QUEUE_FLAG_VM | EXEC_QUEUE_FLAG_MIGRATE));
> +
> job = job_alloc(xe_exec_queue_is_parallel(q) || is_migration);
> if (!job)
> return ERR_PTR(-ENOMEM);
> @@ -119,34 +141,39 @@ struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
> if (err)
> goto err_free;
>
> - for (i = 0; i < q->width; ++i) {
> - struct dma_fence *fence = xe_lrc_alloc_seqno_fence();
> - struct dma_fence_chain *chain;
> -
> - if (IS_ERR(fence)) {
> - err = PTR_ERR(fence);
> - goto err_sched_job;
> + if (!batch_addr) {
> + job->fence = dma_fence_get_stub();
> + job->is_pt_job = true;
> + } else {
> + for (i = 0; i < q->width; ++i) {
> + struct dma_fence *fence = xe_lrc_alloc_seqno_fence();
> + struct dma_fence_chain *chain;
> +
> + if (IS_ERR(fence)) {
> + err = PTR_ERR(fence);
> + goto err_sched_job;
> + }
> + job->ptrs[i].lrc_fence = fence;
> +
> + if (i + 1 == q->width)
> + continue;
> +
> + chain = dma_fence_chain_alloc();
> + if (!chain) {
> + err = -ENOMEM;
> + goto err_sched_job;
> + }
> + job->ptrs[i].chain_fence = chain;
> }
> - job->ptrs[i].lrc_fence = fence;
>
> - if (i + 1 == q->width)
> - continue;
> + width = q->width;
> + if (is_migration)
> + width = 2;
>
> - chain = dma_fence_chain_alloc();
> - if (!chain) {
> - err = -ENOMEM;
> - goto err_sched_job;
> - }
> - job->ptrs[i].chain_fence = chain;
> + for (i = 0; i < width; ++i)
> + job->ptrs[i].batch_addr = batch_addr[i];
> }
>
> - width = q->width;
> - if (is_migration)
> - width = 2;
> -
> - for (i = 0; i < width; ++i)
> - job->ptrs[i].batch_addr = batch_addr[i];
> -
> atomic_inc(&q->job_cnt);
> xe_pm_runtime_get_noresume(job_to_xe(job));
> trace_xe_sched_job_create(job);
> @@ -246,7 +273,7 @@ bool xe_sched_job_completed(struct xe_sched_job *job)
> void xe_sched_job_arm(struct xe_sched_job *job)
> {
> struct xe_exec_queue *q = job->q;
> - struct dma_fence *fence, *prev;
> + struct dma_fence *fence = job->fence, *prev;
> struct xe_vm *vm = q->vm;
> u64 seqno = 0;
> int i;
> @@ -259,6 +286,9 @@ void xe_sched_job_arm(struct xe_sched_job *job)
> xe_vm_assert_held(q->vm);
> }
>
> + if (job->is_pt_job)
> + goto arm;
> +
> if (vm && !xe_sched_job_is_migration(q) && !xe_vm_in_lr_mode(vm) &&
> (vm->batch_invalidate_tlb || vm->tlb_flush_seqno != q->tlb_flush_seqno)) {
> xe_vm_assert_held(vm);
> @@ -286,6 +316,7 @@ void xe_sched_job_arm(struct xe_sched_job *job)
> fence = &chain->base;
> }
>
> +arm:
> job->fence = dma_fence_get(fence); /* Pairs with put in scheduler */
> drm_sched_job_arm(&job->drm);
> }
> @@ -329,7 +360,7 @@ xe_sched_job_snapshot_capture(struct xe_sched_job *job)
>
> snapshot->batch_addr_len = q->width;
> for (i = 0; i < q->width; i++)
> - snapshot->batch_addr[i] =
> + snapshot->batch_addr[i] = job->is_pt_job ? 0 :
> xe_device_uncanonicalize_addr(xe, job->ptrs[i].batch_addr);
>
> return snapshot;
> diff --git a/drivers/gpu/drm/xe/xe_sched_job_types.h b/drivers/gpu/drm/xe/xe_sched_job_types.h
> index 0490b1247a6e..5e1824c36c74 100644
> --- a/drivers/gpu/drm/xe/xe_sched_job_types.h
> +++ b/drivers/gpu/drm/xe/xe_sched_job_types.h
> @@ -10,10 +10,29 @@
>
> #include <drm/gpu_scheduler.h>
>
> -struct xe_exec_queue;
> struct dma_fence;
> struct dma_fence_chain;
>
> +struct xe_exec_queue;
> +struct xe_migrate_pt_update_ops;
> +struct xe_pt_job_ops;
> +struct xe_tile;
> +struct xe_vm;
> +
> +/**
> + * struct xe_pt_update_args - PT update arguments
> + */
> +struct xe_pt_update_args {
> + /** @vm: VM which is being bound */
> + struct xe_vm *vm;
> + /** @tile: Tile which page tables belong to */
> + struct xe_tile *tile;
> + /** @ops: Migrate PT update ops */
> + const struct xe_migrate_pt_update_ops *ops;
> + /** @pt_job_ops: PT job ops state */
> + struct xe_pt_job_ops *pt_job_ops;
> +};
> +
> /**
> * struct xe_job_ptrs - Per hw engine instance data
> */
> @@ -71,8 +90,14 @@ struct xe_sched_job {
> bool restore_replay;
> /** @last_replay: last job being replayed */
> bool last_replay;
> - /** @ptrs: per instance pointers. */
> - struct xe_job_ptrs ptrs[];
> + /** @is_pt_job: is a PT job */
> + bool is_pt_job;
> + union {
> + /** @ptrs: per instance pointers. */
> + DECLARE_FLEX_ARRAY(struct xe_job_ptrs, ptrs);
> + /** @pt_update: PT update arguments */
> + DECLARE_FLEX_ARRAY(struct xe_pt_update_args, pt_update);
> + };
> };
>
> struct xe_sched_job_snapshot {
> diff --git a/drivers/gpu/drm/xe/xe_trace.h b/drivers/gpu/drm/xe/xe_trace.h
> index 2fe8f89a1e34..d4e9d91f6f7f 100644
> --- a/drivers/gpu/drm/xe/xe_trace.h
> +++ b/drivers/gpu/drm/xe/xe_trace.h
> @@ -261,7 +261,7 @@ DECLARE_EVENT_CLASS(xe_sched_job,
> __entry->flags = job->q->flags;
> __entry->error = job->fence ? job->fence->error : 0;
> __entry->fence = job->fence;
> - __entry->batch_addr = (u64)job->ptrs[0].batch_addr;
> + __entry->batch_addr = job->is_pt_job ? 0 : (u64)job->ptrs[0].batch_addr;
> ),
>
> TP_printk("dev=%s, fence=%p, seqno=%u, lrc_seqno=%u, gt=%u, guc_id=%d, batch_addr=0x%012llx, guc_state=0x%x, flags=0x%x, error=%d",
> --
> 2.34.1
>
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 08/24] drm/xe: Add helpers to access PT ops
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (6 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 09/24] drm/xe: Add struct xe_pt_job_ops Matthew Brost
` (20 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Francois Dugast
Add helpers to access PT ops, making it easier to shuffle the location of
the ops structures without requiring widespread code changes.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
---
drivers/gpu/drm/xe/xe_pt.c | 65 ++++++++++++++++++++++++++------------
1 file changed, 45 insertions(+), 20 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index 42a37e40a6c0..2cdfb577e941 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -2088,13 +2088,37 @@ xe_pt_commit_prepare_unbind(struct xe_vma *vma,
}
}
+static struct xe_vm_pgtable_update_op *
+to_pt_op(struct xe_vm_pgtable_update_ops *pt_update_ops, u32 op_idx)
+{
+ return &pt_update_ops->ops[op_idx];
+}
+
+static u32
+get_current_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
+{
+ return pt_update_ops->current_op;
+}
+
+static struct xe_vm_pgtable_update_op *
+to_current_pt_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
+{
+ return to_pt_op(pt_update_ops, get_current_op(pt_update_ops));
+}
+
+static void
+incr_current_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
+{
+ ++pt_update_ops->current_op;
+}
+
static void
xe_pt_update_ops_rfence_interval(struct xe_vm_pgtable_update_ops *pt_update_ops,
u64 start, u64 end)
{
u64 last;
- u32 current_op = pt_update_ops->current_op;
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[current_op];
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_current_pt_op(pt_update_ops);
int i, level = 0;
for (i = 0; i < pt_op->num_entries; i++) {
@@ -2129,8 +2153,8 @@ static int bind_op_prepare(struct xe_vm *vm, struct xe_tile *tile,
struct xe_vm_pgtable_update_ops *pt_update_ops,
struct xe_vma *vma, bool invalidate_on_bind)
{
- u32 current_op = pt_update_ops->current_op;
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[current_op];
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_current_pt_op(pt_update_ops);
int err;
xe_tile_assert(tile, !xe_vma_is_cpu_addr_mirror(vma));
@@ -2159,7 +2183,7 @@ static int bind_op_prepare(struct xe_vm *vm, struct xe_tile *tile,
xe_pt_update_ops_rfence_interval(pt_update_ops,
xe_vma_start(vma),
xe_vma_end(vma));
- ++pt_update_ops->current_op;
+ incr_current_op(pt_update_ops);
pt_update_ops->needs_svm_lock |= xe_vma_is_userptr(vma);
/*
@@ -2202,8 +2226,8 @@ static int bind_range_prepare(struct xe_vm *vm, struct xe_tile *tile,
struct xe_vm_pgtable_update_ops *pt_update_ops,
struct xe_vma *vma, struct xe_svm_range *range)
{
- u32 current_op = pt_update_ops->current_op;
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[current_op];
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_current_pt_op(pt_update_ops);
int err;
xe_tile_assert(tile, xe_vma_is_cpu_addr_mirror(vma));
@@ -2227,7 +2251,7 @@ static int bind_range_prepare(struct xe_vm *vm, struct xe_tile *tile,
xe_pt_update_ops_rfence_interval(pt_update_ops,
xe_svm_range_start(range),
xe_svm_range_end(range));
- ++pt_update_ops->current_op;
+ incr_current_op(pt_update_ops);
pt_update_ops->needs_svm_lock = true;
pt_op->vma = vma;
@@ -2245,8 +2269,8 @@ static int unbind_op_prepare(struct xe_tile *tile,
struct xe_vma *vma)
{
struct xe_device *xe = tile_to_xe(tile);
- u32 current_op = pt_update_ops->current_op;
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[current_op];
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_current_pt_op(pt_update_ops);
int err;
if (!((vma->tile_present | vma->tile_staged) & BIT(tile->id)))
@@ -2285,7 +2309,7 @@ static int unbind_op_prepare(struct xe_tile *tile,
pt_op->num_entries, false);
xe_pt_update_ops_rfence_interval(pt_update_ops, xe_vma_start(vma),
xe_vma_end(vma));
- ++pt_update_ops->current_op;
+ incr_current_op(pt_update_ops);
pt_update_ops->needs_svm_lock |= xe_vma_is_userptr(vma);
pt_update_ops->needs_invalidation = true;
@@ -2325,8 +2349,8 @@ static int unbind_range_prepare(struct xe_vm *vm,
struct xe_vm_pgtable_update_ops *pt_update_ops,
struct xe_svm_range *range)
{
- u32 current_op = pt_update_ops->current_op;
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[current_op];
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_current_pt_op(pt_update_ops);
if (!(range->tile_present & BIT(tile->id)))
return 0;
@@ -2347,7 +2371,7 @@ static int unbind_range_prepare(struct xe_vm *vm,
pt_op->num_entries, false);
xe_pt_update_ops_rfence_interval(pt_update_ops, xe_svm_range_start(range),
xe_svm_range_end(range));
- ++pt_update_ops->current_op;
+ incr_current_op(pt_update_ops);
pt_update_ops->needs_svm_lock = true;
pt_update_ops->needs_invalidation |= xe_vm_has_scratch(vm) ||
xe_vm_has_valid_gpu_mapping(tile, range->tile_present,
@@ -2504,7 +2528,7 @@ int xe_pt_update_ops_prepare(struct xe_tile *tile, struct xe_vma_ops *vops)
return err;
}
- xe_tile_assert(tile, pt_update_ops->current_op <=
+ xe_tile_assert(tile, get_current_op(pt_update_ops) <=
pt_update_ops->num_ops);
#ifdef TEST_VM_OPS_ERROR
@@ -2737,7 +2761,7 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
lockdep_assert_held(&vm->lock);
xe_vm_assert_held(vm);
- if (!pt_update_ops->current_op) {
+ if (!get_current_op(pt_update_ops)) {
xe_tile_assert(tile, xe_vm_in_fault_mode(vm));
return dma_fence_get_stub();
@@ -2805,8 +2829,9 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
}
/* Point of no return - VM killed if failure after this */
- for (i = 0; i < pt_update_ops->current_op; ++i) {
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[i];
+ for (i = 0; i < get_current_op(pt_update_ops); ++i) {
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_pt_op(pt_update_ops, i);
xe_pt_commit(pt_op->vma, pt_op->entries,
pt_op->num_entries, &pt_update_ops->deferred);
@@ -2930,9 +2955,9 @@ void xe_pt_update_ops_abort(struct xe_tile *tile, struct xe_vma_ops *vops)
for (i = pt_update_ops->num_ops - 1; i >= 0; --i) {
struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->ops[i];
+ to_pt_op(pt_update_ops, i);
- if (!pt_op->vma || i >= pt_update_ops->current_op)
+ if (!pt_op->vma || i >= get_current_op(pt_update_ops))
continue;
if (pt_op->bind)
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 09/24] drm/xe: Add struct xe_pt_job_ops
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (7 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 08/24] drm/xe: Add helpers to access PT ops Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
` (19 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Himal Prasad Ghimiray
Add struct xe_pt_job_ops, a dynamically refcounted object that contains
the information required to issue a CPU bind via a job after the initial
bind IOCTL returns.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
v7:
- Drop XE_BO_FLAG_PUT_VM_ASYNC, xe_vm_get and VM ref BO core (Saskio)
---
drivers/gpu/drm/xe/xe_migrate.c | 10 +--
drivers/gpu/drm/xe/xe_pt.c | 126 ++++++++++++++++++++++++++-----
drivers/gpu/drm/xe/xe_pt.h | 4 +
drivers/gpu/drm/xe/xe_pt_types.h | 27 +++++--
drivers/gpu/drm/xe/xe_vm.c | 10 +--
5 files changed, 142 insertions(+), 35 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 40e60567cb58..d7f35aa91c1c 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -1873,7 +1873,7 @@ xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
}
xe_migrate_update_pgtables_cpu_execute(vm, m->tile, ops,
- pt_update_ops->ops,
+ pt_update_ops->pt_job_ops->ops,
pt_update_ops->num_ops);
return dma_fence_get_stub();
@@ -1900,7 +1900,7 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
bool usm = is_migrate && xe->info.has_usm;
for (i = 0; i < pt_update_ops->num_ops; ++i) {
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[i];
+ struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->pt_job_ops->ops[i];
struct xe_vm_pgtable_update *updates = pt_op->entries;
num_updates += pt_op->num_entries;
@@ -1969,7 +1969,7 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
for (; i < pt_update_ops->num_ops; ++i) {
struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->ops[i];
+ &pt_update_ops->pt_job_ops->ops[i];
struct xe_vm_pgtable_update *updates = pt_op->entries;
for (; j < pt_op->num_entries; ++j, ++current_update, ++idx) {
@@ -2006,7 +2006,7 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
(page_ofs / sizeof(u64)) * XE_PAGE_SIZE;
for (i = 0; i < pt_update_ops->num_ops; ++i) {
struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->ops[i];
+ &pt_update_ops->pt_job_ops->ops[i];
struct xe_vm_pgtable_update *updates = pt_op->entries;
for (j = 0; j < pt_op->num_entries; ++j) {
@@ -2024,7 +2024,7 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
for (i = 0; i < pt_update_ops->num_ops; ++i) {
struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->ops[i];
+ &pt_update_ops->pt_job_ops->ops[i];
struct xe_vm_pgtable_update *updates = pt_op->entries;
for (j = 0; j < pt_op->num_entries; ++j)
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index 2cdfb577e941..c004ce449b69 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -206,6 +206,7 @@ unsigned int xe_pt_shift(unsigned int level)
*/
void xe_pt_destroy(struct xe_pt *pt, u32 flags, struct llist_head *deferred)
{
+ bool added = false;
int i;
if (!pt)
@@ -213,7 +214,9 @@ void xe_pt_destroy(struct xe_pt *pt, u32 flags, struct llist_head *deferred)
XE_WARN_ON(!list_empty(&pt->bo->ttm.base.gpuva.list));
xe_bo_unpin(pt->bo);
- xe_bo_put_deferred(pt->bo, deferred, NULL);
+ xe_bo_put_deferred(pt->bo, deferred, &added);
+ if (added)
+ xe_assert(pt->bo->vm->xe, !kref_read(&pt->bo->ttm.base.refcount));
if (pt->level > 0 && pt->num_live) {
struct xe_pt_dir *pt_dir = as_xe_pt_dir(pt);
@@ -2091,13 +2094,13 @@ xe_pt_commit_prepare_unbind(struct xe_vma *vma,
static struct xe_vm_pgtable_update_op *
to_pt_op(struct xe_vm_pgtable_update_ops *pt_update_ops, u32 op_idx)
{
- return &pt_update_ops->ops[op_idx];
+ return &pt_update_ops->pt_job_ops->ops[op_idx];
}
static u32
get_current_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
{
- return pt_update_ops->current_op;
+ return pt_update_ops->pt_job_ops->current_op;
}
static struct xe_vm_pgtable_update_op *
@@ -2109,7 +2112,7 @@ to_current_pt_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
static void
incr_current_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
{
- ++pt_update_ops->current_op;
+ ++pt_update_ops->pt_job_ops->current_op;
}
static void
@@ -2483,8 +2486,7 @@ static int op_prepare(struct xe_vm *vm,
static void
xe_pt_update_ops_init(struct xe_vm_pgtable_update_ops *pt_update_ops)
{
- init_llist_head(&pt_update_ops->deferred);
- pt_update_ops->current_op = 0;
+ pt_update_ops->pt_job_ops->current_op = 0;
pt_update_ops->start = ~0x0ull;
pt_update_ops->last = 0x0ull;
pt_update_ops->needs_svm_lock = false;
@@ -2834,7 +2836,8 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
to_pt_op(pt_update_ops, i);
xe_pt_commit(pt_op->vma, pt_op->entries,
- pt_op->num_entries, &pt_update_ops->deferred);
+ pt_op->num_entries,
+ &pt_update_ops->pt_job_ops->deferred);
pt_op->vma = NULL; /* skip in xe_pt_update_ops_abort */
}
@@ -2922,19 +2925,8 @@ void xe_pt_update_ops_fini(struct xe_tile *tile, struct xe_vma_ops *vops)
{
struct xe_vm_pgtable_update_ops *pt_update_ops =
&vops->pt_update_ops[tile->id];
- int i;
xe_page_reclaim_entries_put(pt_update_ops->prl.entries);
-
- lockdep_assert_held(&vops->vm->lock);
- xe_vm_assert_held(vops->vm);
-
- for (i = 0; i < pt_update_ops->current_op; ++i) {
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->ops[i];
-
- xe_pt_free_bind(pt_op->entries, pt_op->num_entries);
- }
- xe_bo_put_commit(&vops->pt_update_ops[tile->id].deferred);
}
/**
@@ -2971,3 +2963,101 @@ void xe_pt_update_ops_abort(struct xe_tile *tile, struct xe_vma_ops *vops)
xe_pt_update_ops_fini(tile, vops);
}
+
+/**
+ * xe_pt_job_ops_alloc() - Allocate PT job ops
+ * @num_ops: Number of VM PT update ops
+ *
+ * Allocate PT job ops and internal array of VM PT update ops.
+ *
+ * Return: Pointer to PT job ops or NULL
+ */
+struct xe_pt_job_ops *xe_pt_job_ops_alloc(u32 num_ops)
+{
+ struct xe_pt_job_ops *pt_job_ops;
+
+ pt_job_ops = kmalloc_obj(*pt_job_ops);
+ if (!pt_job_ops)
+ return NULL;
+
+ pt_job_ops->ops = kvmalloc_array(num_ops, sizeof(*pt_job_ops->ops),
+ GFP_KERNEL);
+ if (!pt_job_ops->ops) {
+ kfree(pt_job_ops);
+ return NULL;
+ }
+
+ pt_job_ops->current_op = 0;
+ kref_init(&pt_job_ops->refcount);
+ init_llist_head(&pt_job_ops->deferred);
+
+ return pt_job_ops;
+}
+
+/**
+ * xe_pt_job_ops_get() - Get PT job ops
+ * @pt_job_ops: PT job ops to get
+ *
+ * Take a reference to PT job ops
+ *
+ * Return: Pointer to PT job ops or NULL
+ */
+struct xe_pt_job_ops *xe_pt_job_ops_get(struct xe_pt_job_ops *pt_job_ops)
+{
+ if (pt_job_ops)
+ kref_get(&pt_job_ops->refcount);
+
+ return pt_job_ops;
+}
+
+static void xe_pt_update_ops_free(struct xe_vm_pgtable_update_op *pt_op,
+ u32 num_ops)
+{
+ u32 i;
+
+ for (i = 0; i < num_ops; ++i, ++pt_op)
+ xe_pt_free_bind(pt_op->entries, pt_op->num_entries);
+}
+
+static void xe_pt_job_ops_destroy(struct kref *ref)
+{
+ struct xe_pt_job_ops *pt_job_ops =
+ container_of(ref, struct xe_pt_job_ops, refcount);
+ struct llist_node *freed;
+ struct xe_bo *bo, *next;
+
+ xe_pt_update_ops_free(pt_job_ops->ops,
+ pt_job_ops->current_op);
+
+ freed = llist_del_all(&pt_job_ops->deferred);
+ if (freed) {
+ llist_for_each_entry_safe(bo, next, freed, freed) {
+ struct xe_bo_dev *bo_device =
+ &xe_bo_device(bo)->bo_device;
+ /*
+ * If called from run_job, we are in the dma-fencing
+ * path and cannot take dma-resv locks so use an async
+ * put.
+ */
+ if (llist_add(&bo->freed, &bo_device->async_list))
+ schedule_work(&bo_device->async_free);
+ }
+ }
+
+ kvfree(pt_job_ops->ops);
+ kfree(pt_job_ops);
+}
+
+/**
+ * xe_pt_job_ops_put() - Put PT job ops
+ * @pt_job_ops: PT job ops to put
+ *
+ * Drop a reference to PT job ops
+ */
+void xe_pt_job_ops_put(struct xe_pt_job_ops *pt_job_ops)
+{
+ if (!pt_job_ops)
+ return;
+
+ kref_put(&pt_job_ops->refcount, xe_pt_job_ops_destroy);
+}
diff --git a/drivers/gpu/drm/xe/xe_pt.h b/drivers/gpu/drm/xe/xe_pt.h
index 4daeebaab5a1..5faddb8e700c 100644
--- a/drivers/gpu/drm/xe/xe_pt.h
+++ b/drivers/gpu/drm/xe/xe_pt.h
@@ -49,4 +49,8 @@ bool xe_pt_zap_ptes(struct xe_tile *tile, struct xe_vma *vma);
bool xe_pt_zap_ptes_range(struct xe_tile *tile, struct xe_vm *vm,
struct xe_svm_range *range);
+struct xe_pt_job_ops *xe_pt_job_ops_alloc(u32 num_ops);
+struct xe_pt_job_ops *xe_pt_job_ops_get(struct xe_pt_job_ops *pt_job_ops);
+void xe_pt_job_ops_put(struct xe_pt_job_ops *pt_job_ops);
+
#endif
diff --git a/drivers/gpu/drm/xe/xe_pt_types.h b/drivers/gpu/drm/xe/xe_pt_types.h
index a7d1bb708b69..39c5b89ce9b7 100644
--- a/drivers/gpu/drm/xe/xe_pt_types.h
+++ b/drivers/gpu/drm/xe/xe_pt_types.h
@@ -91,12 +91,29 @@ struct xe_vm_pgtable_update_op {
bool rebind;
};
+/**
+ * struct xe_pt_job_ops - Page-table update operations (dynamically allocated)
+ *
+ * This is the portion of &struct xe_vma_ops and
+ * &struct xe_vm_pgtable_update_ops that is dynamically allocated, as it
+ * must remain valid until the associated bind job completes. A reference
+ * count controls its lifetime.
+ */
+struct xe_pt_job_ops {
+ /** @current_op: current page-table update operation */
+ u32 current_op;
+ /** @refcount: reference count */
+ struct kref refcount;
+ /** @deferred: list of deferred PT entries to destroy */
+ struct llist_head deferred;
+ /** @ops: page-table update operations */
+ struct xe_vm_pgtable_update_op *ops;
+};
+
/** struct xe_vm_pgtable_update_ops: page table update operations */
struct xe_vm_pgtable_update_ops {
- /** @ops: operations */
- struct xe_vm_pgtable_update_op *ops;
- /** @deferred: deferred list to destroy PT entries */
- struct llist_head deferred;
+ /** @pt_job_ops: PT update operations dynamic allocation*/
+ struct xe_pt_job_ops *pt_job_ops;
/** @q: exec queue for PT operations */
struct xe_exec_queue *q;
/** @prl: embedded page reclaim list */
@@ -107,8 +124,6 @@ struct xe_vm_pgtable_update_ops {
u64 last;
/** @num_ops: number of operations */
u32 num_ops;
- /** @current_op: current operations */
- u32 current_op;
/** @needs_svm_lock: Needs SVM lock */
bool needs_svm_lock;
/** @needs_invalidation: Needs invalidation */
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index 390da884c727..a441c5a2a8c9 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -682,11 +682,9 @@ static int xe_vma_ops_alloc(struct xe_vma_ops *vops, bool array_of_binds)
if (!vops->pt_update_ops[i].num_ops)
continue;
- vops->pt_update_ops[i].ops =
- kmalloc_objs(*vops->pt_update_ops[i].ops,
- vops->pt_update_ops[i].num_ops,
- GFP_KERNEL | __GFP_RETRY_MAYFAIL | __GFP_NOWARN);
- if (!vops->pt_update_ops[i].ops)
+ vops->pt_update_ops[i].pt_job_ops =
+ xe_pt_job_ops_alloc(vops->pt_update_ops[i].num_ops);
+ if (!vops->pt_update_ops[i].pt_job_ops)
return array_of_binds ? -ENOBUFS : -ENOMEM;
}
@@ -733,7 +731,7 @@ static void xe_vma_ops_fini(struct xe_vma_ops *vops)
xe_vma_svm_prefetch_ops_fini(vops);
for (i = 0; i < XE_MAX_TILES_PER_DEVICE; ++i)
- kfree(vops->pt_update_ops[i].ops);
+ xe_pt_job_ops_put(vops->pt_update_ops[i].pt_job_ops);
}
static void xe_vma_ops_incr_pt_update_ops(struct xe_vma_ops *vops, u8 tile_mask, int inc_val)
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (8 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 09/24] drm/xe: Add struct xe_pt_job_ops Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 5:39 ` sashiko-bot
2026-09-25 4:53 ` [PATCH v7 11/24] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
` (18 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe
PT jobs bypass GPU execution for the final step of a bind job, using the
CPU to program the required page tables. Teach the GuC submission backend
how to execute these jobs.
PT job submission is implemented in the GuC backend for simplicity. A
follow-up patch could introduce a dedicated backend for PT jobs.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
---
v7:
- Also run PT jobs, for cancelled force clear (Himal)
---
drivers/gpu/drm/xe/xe_guc_submit.c | 31 +++++++++++++++++++++++++++---
drivers/gpu/drm/xe/xe_migrate.c | 20 +++++++++++++++----
drivers/gpu/drm/xe/xe_migrate.h | 7 +++++++
3 files changed, 51 insertions(+), 7 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index 4bd1ead57efa..9faddb7c6407 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -39,9 +39,11 @@
#include "xe_lrc.h"
#include "xe_macros.h"
#include "xe_map.h"
+#include "xe_migrate.h"
#include "xe_mocs.h"
#include "xe_module.h"
#include "xe_pm.h"
+#include "xe_pt.h"
#include "xe_ring_ops_types.h"
#include "xe_sched_job.h"
#include "xe_sleep.h"
@@ -1238,21 +1240,44 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
}
}
+static bool is_pt_job(struct xe_sched_job *job)
+{
+ return job->is_pt_job;
+}
+
+static void run_pt_job(struct xe_sched_job *job, bool force_clear)
+{
+ xe_migrate_update_pgtables_cpu_execute(job->pt_update[0].vm,
+ job->pt_update[0].tile,
+ job->pt_update[0].ops,
+ job->pt_update[0].pt_job_ops->ops,
+ job->pt_update[0].pt_job_ops->current_op,
+ force_clear);
+}
+
static struct dma_fence *
guc_exec_queue_run_job(struct drm_sched_job *drm_job)
{
struct xe_sched_job *job = to_xe_sched_job(drm_job);
struct xe_exec_queue *q = job->q;
struct xe_guc *guc = exec_queue_to_guc(q);
- bool killed_or_banned_or_wedged =
- exec_queue_killed_or_banned_or_wedged(q);
+ bool killed_or_banned_or_wedged_or_error =
+ exec_queue_killed_or_banned_or_wedged(q) ||
+ xe_sched_job_is_error(job);
xe_gt_assert(guc_to_gt(guc), !(exec_queue_destroyed(q) || exec_queue_pending_disable(q)) ||
exec_queue_banned(q) || exec_queue_suspended(q));
trace_xe_sched_job_run(job);
- if (!killed_or_banned_or_wedged && !xe_sched_job_is_error(job)) {
+ if (is_pt_job(job)) {
+ xe_gt_assert(guc_to_gt(guc), !exec_queue_registered(q));
+ run_pt_job(job, killed_or_banned_or_wedged_or_error);
+ xe_pt_job_ops_put(job->pt_update[0].pt_job_ops);
+ dma_fence_put(job->fence); /* Drop ref from xe_sched_job_arm */
+
+ return NULL;
+ } else if (!killed_or_banned_or_wedged_or_error) {
if (xe_exec_queue_is_multi_queue_secondary(q)) {
struct xe_exec_queue *primary = xe_exec_queue_multi_queue_primary(q);
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index d7f35aa91c1c..3007fe0b8535 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -1819,11 +1819,23 @@ struct migrate_test_params {
container_of(_priv, struct migrate_test_params, base)
#endif
-static void
+/**
+ * xe_migrate_update_pgtables_cpu_execute() - Update a VM's PTEs via the CPU
+ * @vm: The VM being updated
+ * @tile: The tile being updated
+ * @ops: The migrate PT update ops
+ * @pt_ops: The VM PT update ops
+ * @num_ops: The number of The VM PT update ops
+ * @force_clear: Force clear
+ *
+ * Execute the VM PT update ops array which results in a VM's PTEs being updated
+ * via the CPU.
+ */
+void
xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
const struct xe_migrate_pt_update_ops *ops,
struct xe_vm_pgtable_update_op *pt_op,
- u32 num_ops)
+ u32 num_ops, bool force_clear)
{
u32 j, i;
@@ -1834,7 +1846,7 @@ xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
xe_tile_assert(tile, !iosys_map_is_null(&update->pt_bo->vmap));
- if (pt_op->bind)
+ if (pt_op->bind && !force_clear)
ops->populate(tile, &update->pt_bo->vmap,
NULL, update->ofs, update->qwords,
update);
@@ -1874,7 +1886,7 @@ xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
xe_migrate_update_pgtables_cpu_execute(vm, m->tile, ops,
pt_update_ops->pt_job_ops->ops,
- pt_update_ops->num_ops);
+ pt_update_ops->num_ops, false);
return dma_fence_get_stub();
}
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index 2ec9de896dfe..c41f6e3ca2c6 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -24,6 +24,7 @@ struct xe_pt;
struct xe_tile;
struct xe_vm;
struct xe_vm_pgtable_update;
+struct xe_vm_pgtable_update_op;
struct xe_vma;
enum xe_sriov_vf_ccs_rw_ctxs;
@@ -163,6 +164,12 @@ struct dma_fence *xe_migrate_clear(struct xe_migrate *m,
struct xe_vm *xe_migrate_get_vm(struct xe_migrate *m);
+void
+xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
+ const struct xe_migrate_pt_update_ops *ops,
+ struct xe_vm_pgtable_update_op *pt_op,
+ u32 num_ops, bool force_clear);
+
struct dma_fence *
xe_migrate_update_pgtables(struct xe_migrate *m,
struct xe_migrate_pt_update *pt_update);
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs
2026-09-25 4:53 ` [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
@ 2026-09-25 5:39 ` sashiko-bot
2026-09-25 5:58 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 5:39 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] xe: memory leak of pt_job_ops when PT job is aborted before execution
- [High] xe: PT jobs swallow cancellation error statuses and signal false success
--
commit 784a2cd005e51d36d34825eebaf5abff98891eff
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: Update GuC submission backend to run PT jobs
PT jobs bypass GPU execution for the final step of a bind job, using the
CPU to program the required page tables. Teach the GuC submission backend
how to execute these jobs.
> diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> index 4bd1ead57efa4..9faddb7c6407f 100644
> --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
[ ... ]
> @@ -1238,21 +1240,44 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
> static struct dma_fence *
> guc_exec_queue_run_job(struct drm_sched_job *drm_job)
> {
> struct xe_sched_job *job = to_xe_sched_job(drm_job);
> struct xe_exec_queue *q = job->q;
> struct xe_guc *guc = exec_queue_to_guc(q);
> - bool killed_or_banned_or_wedged =
> - exec_queue_killed_or_banned_or_wedged(q);
> + bool killed_or_banned_or_wedged_or_error =
> + exec_queue_killed_or_banned_or_wedged(q) ||
> + xe_sched_job_is_error(job);
[Severity: High]
Can this condition fail to detect errors for PT jobs?
PT jobs are initialized with a global stub fence. When the queue is
killed or banned, xe_sched_job_set_error() attempts to set an error on this
fence but fails because the stub fence is already permanently signaled. As a
result, xe_sched_job_is_error() will always be false for PT jobs.
> xe_gt_assert(guc_to_gt(guc), !(exec_queue_destroyed(q) || exec_queue_pending_disable(q)) ||
> exec_queue_banned(q) || exec_queue_suspended(q));
>
> trace_xe_sched_job_run(job);
>
> - if (!killed_or_banned_or_wedged && !xe_sched_job_is_error(job)) {
> + if (is_pt_job(job)) {
> + xe_gt_assert(guc_to_gt(guc), !exec_queue_registered(q));
> + run_pt_job(job, killed_or_banned_or_wedged_or_error);
> + xe_pt_job_ops_put(job->pt_update[0].pt_job_ops);
[Severity: High]
Does this leak the pt_job_ops when a PT job is aborted before execution?
If the DRM scheduler is stopped, the entity killed, or the execution queue
banned before run_job is invoked, the scheduler calls free_job directly
without running it. The free_job callback invokes xe_sched_job_destroy(),
which then calls xe_sched_job_free_fences(). That function explicitly skips
PT jobs via an early return:
xe_sched_job_free_fences() {
if (job->is_pt_job)
return;
...
}
Since xe_pt_job_ops_put() is only called here in guc_exec_queue_run_job(),
which is completely bypassed on this cancellation path, does the refcount
leak?
> + dma_fence_put(job->fence); /* Drop ref from xe_sched_job_arm */
> +
> + return NULL;
[Severity: High]
Does this swallow cancellation error statuses and signal false success for
PT jobs?
Even when killed_or_banned_or_wedged_or_error is true, this path
unconditionally returns NULL. The DRM scheduler treats a NULL return from
run_job as an immediate successful completion, signaling the software
finished fence with 0 instead of an error (like -ECANCELED).
Userspace waiting on a failed or cancelled bind will observe false success
and might proceed to access memory that hasn't been properly mapped.
> + } else if (!killed_or_banned_or_wedged_or_error) {
> if (xe_exec_queue_is_multi_queue_secondary(q)) {
> struct xe_exec_queue *primary = xe_exec_queue_multi_queue_primary(q);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=10
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs
2026-09-25 5:39 ` sashiko-bot
@ 2026-09-25 5:58 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 5:58 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 05:39:14AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [High] xe: memory leak of pt_job_ops when PT job is aborted before execution
> - [High] xe: PT jobs swallow cancellation error statuses and signal false success
> --
>
> commit 784a2cd005e51d36d34825eebaf5abff98891eff
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Update GuC submission backend to run PT jobs
>
> PT jobs bypass GPU execution for the final step of a bind job, using the
> CPU to program the required page tables. Teach the GuC submission backend
> how to execute these jobs.
>
> > diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> > index 4bd1ead57efa4..9faddb7c6407f 100644
> > --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> > +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
> [ ... ]
> > @@ -1238,21 +1240,44 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
> > static struct dma_fence *
> > guc_exec_queue_run_job(struct drm_sched_job *drm_job)
> > {
> > struct xe_sched_job *job = to_xe_sched_job(drm_job);
> > struct xe_exec_queue *q = job->q;
> > struct xe_guc *guc = exec_queue_to_guc(q);
> > - bool killed_or_banned_or_wedged =
> > - exec_queue_killed_or_banned_or_wedged(q);
> > + bool killed_or_banned_or_wedged_or_error =
> > + exec_queue_killed_or_banned_or_wedged(q) ||
> > + xe_sched_job_is_error(job);
>
> [Severity: High]
> Can this condition fail to detect errors for PT jobs?
>
> PT jobs are initialized with a global stub fence. When the queue is
> killed or banned, xe_sched_job_set_error() attempts to set an error on this
> fence but fails because the stub fence is already permanently signaled. As a
> result, xe_sched_job_is_error() will always be false for PT jobs.
>
Non-issue - the TDR actually should never run for PT jobs, when split PT
jobs into its own backend (follow up) that will be clear, but sematic of
PT jobs is we always run them but if killed, bannded, or wedged we issue
a clear. This sematic holds.
> > xe_gt_assert(guc_to_gt(guc), !(exec_queue_destroyed(q) || exec_queue_pending_disable(q)) ||
> > exec_queue_banned(q) || exec_queue_suspended(q));
> >
> > trace_xe_sched_job_run(job);
> >
> > - if (!killed_or_banned_or_wedged && !xe_sched_job_is_error(job)) {
> > + if (is_pt_job(job)) {
> > + xe_gt_assert(guc_to_gt(guc), !exec_queue_registered(q));
> > + run_pt_job(job, killed_or_banned_or_wedged_or_error);
> > + xe_pt_job_ops_put(job->pt_update[0].pt_job_ops);
>
> [Severity: High]
> Does this leak the pt_job_ops when a PT job is aborted before execution?
>
A call _arm() guarantees a call of run_job(), thus xe_pt_job_ops_put is
always called. Non-issue.
> If the DRM scheduler is stopped, the entity killed, or the execution queue
> banned before run_job is invoked, the scheduler calls free_job directly
> without running it. The free_job callback invokes xe_sched_job_destroy(),
> which then calls xe_sched_job_free_fences(). That function explicitly skips
> PT jobs via an early return:
>
> xe_sched_job_free_fences() {
> if (job->is_pt_job)
> return;
> ...
> }
>
> Since xe_pt_job_ops_put() is only called here in guc_exec_queue_run_job(),
> which is completely bypassed on this cancellation path, does the refcount
> leak?
>
> > + dma_fence_put(job->fence); /* Drop ref from xe_sched_job_arm */
> > +
> > + return NULL;
>
> [Severity: High]
> Does this swallow cancellation error statuses and signal false success for
> PT jobs?
>
> Even when killed_or_banned_or_wedged_or_error is true, this path
> unconditionally returns NULL. The DRM scheduler treats a NULL return from
> run_job as an immediate successful completion, signaling the software
> finished fence with 0 instead of an error (like -ECANCELED).
>
> Userspace waiting on a failed or cancelled bind will observe false success
> and might proceed to access memory that hasn't been properly mapped.
>
I'd say non-issue as AFIAK user space never latches onto as fence state.
Matt
> > + } else if (!killed_or_banned_or_wedged_or_error) {
> > if (xe_exec_queue_is_multi_queue_secondary(q)) {
> > struct xe_exec_queue *primary = xe_exec_queue_multi_queue_primary(q);
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=10
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 11/24] drm/xe: Store level in struct xe_vm_pgtable_update
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (9 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 12/24] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
` (17 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Stuart Summers, Himal Prasad Ghimiray
The level was previously extracted from struct xe_pt inside
xe_vm_pgtable_update during CPU binds, which always occurred during the
bind IOCTL. With CPU binds now supported in bind jobs, struct xe_pt may
no longer be valid in memory at that point. To address this, store the
level directly in struct xe_vm_pgtable_update.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Stuart Summers <stuart.summers@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
drivers/gpu/drm/xe/xe_pt.c | 3 ++-
drivers/gpu/drm/xe/xe_pt_types.h | 8 +++++++-
2 files changed, 9 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index c004ce449b69..a483158f94d4 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -380,6 +380,7 @@ xe_pt_new_shared(struct xe_walk_update *wupd, struct xe_pt *parent,
entry->flags = 0;
entry->qwords = 0;
entry->pt_bo->update_index = -1;
+ entry->level = parent->level;
if (alloc_entries) {
entry->pt_entries = kmalloc_objs(*entry->pt_entries, XE_PDES);
@@ -2025,7 +2026,7 @@ xe_migrate_clear_pgtable_callback(struct xe_vm *vm, struct xe_tile *tile,
u32 qword_ofs, u32 num_qwords,
const struct xe_vm_pgtable_update *update)
{
- u64 empty = __xe_pt_empty_pte(tile, vm, update->pt->level);
+ u64 empty = __xe_pt_empty_pte(tile, vm, update->level);
int i;
if (map && map->is_iomem)
diff --git a/drivers/gpu/drm/xe/xe_pt_types.h b/drivers/gpu/drm/xe/xe_pt_types.h
index 39c5b89ce9b7..ccab6613385f 100644
--- a/drivers/gpu/drm/xe/xe_pt_types.h
+++ b/drivers/gpu/drm/xe/xe_pt_types.h
@@ -65,12 +65,18 @@ struct xe_vm_pgtable_update {
/** @qwords: number of PTE's to write */
u32 qwords;
- /** @pt: opaque pointer useful for the caller of xe_migrate_update_pgtables */
+ /**
+ * @pt: opaque pointer useful for PT building in the bind IOCTL. Only
+ * safe to touch during the bind IOCTL (i.e., not in bind jobs).
+ */
struct xe_pt *pt;
/** @pt_entries: Newly added pagetable entries */
struct xe_pt_entry *pt_entries;
+ /** @level: level of update */
+ unsigned int level;
+
/** @flags: Target flags */
u32 flags;
};
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 12/24] drm/xe: Don't use migrate exec queue for page fault binds
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (10 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 11/24] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
` (16 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Himal Prasad Ghimiray
Now that the CPU is always used for binds even in jobs, CPU bind jobs
can pass GPU jobs in the same exec queue resulting dma-fences signaling
out-of-order. Use a dedicated exec queue for binds issued from page
faults to avoid ordering issues and avoid blocking kernel binds on
unrelated copies / clears.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
drivers/gpu/drm/xe/xe_migrate.c | 47 ++++++++++++++++++++++++++++++---
drivers/gpu/drm/xe/xe_migrate.h | 1 +
drivers/gpu/drm/xe/xe_vm.c | 17 +++++++-----
3 files changed, 55 insertions(+), 10 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 3007fe0b8535..d8d35671a190 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -51,6 +51,8 @@
struct xe_migrate {
/** @q: Default exec queue used for migration */
struct xe_exec_queue *q;
+ /** @bind_q: Default exec queue used for binds */
+ struct xe_exec_queue *bind_q;
/** @tile: Backpointer to the tile this struct xe_migrate belongs to. */
struct xe_tile *tile;
/** @job_mutex: Timeline mutex for @eng. */
@@ -115,6 +117,7 @@ static void xe_migrate_fini(void *arg)
mutex_destroy(&m->job_mutex);
xe_vm_close_and_put(m->q->vm);
xe_exec_queue_put(m->q);
+ xe_exec_queue_put(m->bind_q);
}
static inline u16 xe_migrate_pat_index(struct xe_device *xe,
@@ -504,6 +507,15 @@ int xe_migrate_init(struct xe_migrate *m)
goto err_out;
}
+ m->bind_q = xe_exec_queue_create(xe, vm, logical_mask, 1, hwe0,
+ EXEC_QUEUE_FLAG_KERNEL |
+ EXEC_QUEUE_FLAG_HIGH_PRIORITY |
+ EXEC_QUEUE_FLAG_MIGRATE, 0);
+ if (IS_ERR(m->bind_q)) {
+ err = PTR_ERR(m->bind_q);
+ goto err_out;
+ }
+
/*
* XXX: Currently only reserving 1 (likely slow) BCS instance on
* PVC, may want to revisit if performance is needed.
@@ -514,6 +526,15 @@ int xe_migrate_init(struct xe_migrate *m)
EXEC_QUEUE_FLAG_MIGRATE |
EXEC_QUEUE_FLAG_LOW_LATENCY, 0);
} else {
+ m->bind_q = xe_exec_queue_create_class(xe, primary_gt, vm,
+ XE_ENGINE_CLASS_COPY,
+ EXEC_QUEUE_FLAG_KERNEL |
+ EXEC_QUEUE_FLAG_MIGRATE, 0);
+ if (IS_ERR(m->bind_q)) {
+ err = PTR_ERR(m->bind_q);
+ goto err_out;
+ }
+
m->q = xe_exec_queue_create_class(xe, primary_gt, vm,
XE_ENGINE_CLASS_COPY,
EXEC_QUEUE_FLAG_KERNEL |
@@ -549,6 +570,8 @@ int xe_migrate_init(struct xe_migrate *m)
return err;
err_out:
+ if (!IS_ERR_OR_NULL(m->bind_q))
+ xe_exec_queue_put(m->bind_q);
xe_vm_close_and_put(vm);
return err;
@@ -1505,6 +1528,17 @@ static u32 blt_mem_set_cmd_len(struct xe_device *xe)
return 7;
}
+/**
+ * xe_get_migrate_bind_queue() - Get the bind queue from migrate context.
+ * @migrate: Migrate context.
+ *
+ * Return: Pointer to bind queue on success, error on failure
+ */
+struct xe_exec_queue *xe_migrate_bind_queue(struct xe_migrate *migrate)
+{
+ return migrate->bind_q;
+}
+
static void emit_clear_link_copy(struct xe_gt *gt, struct xe_bb *bb, u64 src_ofs,
u32 size, u32 pitch)
{
@@ -1891,6 +1925,11 @@ xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
return dma_fence_get_stub();
}
+static bool is_migrate_queue(struct xe_migrate *m, struct xe_exec_queue *q)
+{
+ return m->bind_q == q;
+}
+
static struct dma_fence *
__xe_migrate_update_pgtables(struct xe_migrate *m,
struct xe_migrate_pt_update *pt_update,
@@ -1908,7 +1947,7 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
u32 num_updates = 0, current_update = 0;
u64 addr;
int err = 0;
- bool is_migrate = pt_update_ops->q == m->q;
+ bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
bool usm = is_migrate && xe->info.has_usm;
for (i = 0; i < pt_update_ops->num_ops; ++i) {
@@ -2630,7 +2669,7 @@ int xe_migrate_access_memory(struct xe_migrate *m, struct xe_bo *bo,
*/
void xe_migrate_job_lock(struct xe_migrate *m, struct xe_exec_queue *q)
{
- bool is_migrate = q == m->q;
+ bool is_migrate = is_migrate_queue(m, q);
if (is_migrate)
mutex_lock(&m->job_mutex);
@@ -2648,7 +2687,7 @@ void xe_migrate_job_lock(struct xe_migrate *m, struct xe_exec_queue *q)
*/
void xe_migrate_job_unlock(struct xe_migrate *m, struct xe_exec_queue *q)
{
- bool is_migrate = q == m->q;
+ bool is_migrate = is_migrate_queue(m, q);
if (is_migrate)
mutex_unlock(&m->job_mutex);
@@ -2665,7 +2704,7 @@ void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
{
struct xe_migrate *m = gt_to_tile(q->gt)->migrate;
- xe_gt_assert(q->gt, q == m->q);
+ xe_gt_assert(q->gt, q == m->bind_q);
lockdep_assert_held(&m->job_mutex);
}
#endif
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index c41f6e3ca2c6..e6b7070134e1 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -146,6 +146,7 @@ void xe_migrate_ccs_rw_copy_clear(struct xe_bo *src_bo,
struct xe_lrc *xe_migrate_lrc(struct xe_migrate *migrate);
struct xe_exec_queue *xe_migrate_exec_queue(struct xe_migrate *migrate);
+struct xe_exec_queue *xe_migrate_bind_queue(struct xe_migrate *migrate);
struct dma_fence *xe_migrate_vram_copy_chunk(struct xe_bo *vram_bo, u64 vram_offset,
struct xe_bo *sysmem_bo, u64 sysmem_offset,
u64 size, enum xe_migrate_copy_dir dir);
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index a441c5a2a8c9..c0d8c0e266a3 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -796,7 +796,9 @@ int xe_vm_rebind(struct xe_vm *vm, bool rebind_worker)
struct xe_vma *vma, *next;
struct xe_vma_ops vops;
struct xe_vma_op *op, *next_op;
- int err, i;
+ struct xe_tile *tile;
+ u8 id;
+ int err;
lockdep_assert_held(&vm->lock);
if ((xe_vm_in_lr_mode(vm) && !rebind_worker) ||
@@ -804,8 +806,11 @@ int xe_vm_rebind(struct xe_vm *vm, bool rebind_worker)
return 0;
xe_vma_ops_init(&vops, vm, NULL, NULL, 0);
- for (i = 0; i < XE_MAX_TILES_PER_DEVICE; ++i)
- vops.pt_update_ops[i].wait_vm_bookkeep = true;
+ for_each_tile(tile, vm->xe, id) {
+ vops.pt_update_ops[id].wait_vm_bookkeep = true;
+ vops.pt_update_ops[id].q =
+ xe_migrate_bind_queue(tile->migrate);
+ }
xe_vm_assert_held(vm);
list_for_each_entry(vma, &vm->rebind_list, combined_links.rebind) {
@@ -863,7 +868,7 @@ struct dma_fence *xe_vma_rebind(struct xe_vm *vm, struct xe_vma *vma, u8 tile_ma
for_each_tile(tile, vm->xe, id) {
vops.pt_update_ops[id].wait_vm_bookkeep = true;
vops.pt_update_ops[tile->id].q =
- xe_migrate_exec_queue(tile->migrate);
+ xe_migrate_bind_queue(tile->migrate);
}
err = xe_vm_ops_add_rebind(&vops, vma, tile_mask);
@@ -955,7 +960,7 @@ struct dma_fence *xe_vm_range_rebind(struct xe_vm *vm,
for_each_tile(tile, vm->xe, id) {
vops.pt_update_ops[id].wait_vm_bookkeep = true;
vops.pt_update_ops[tile->id].q =
- xe_migrate_exec_queue(tile->migrate);
+ xe_migrate_bind_queue(tile->migrate);
}
err = xe_vm_ops_add_range_rebind(&vops, vma, range, tile_mask);
@@ -1039,7 +1044,7 @@ struct dma_fence *xe_vm_range_unbind(struct xe_vm *vm,
for_each_tile(tile, vm->xe, id) {
vops.pt_update_ops[id].wait_vm_bookkeep = true;
vops.pt_update_ops[tile->id].q =
- xe_migrate_exec_queue(tile->migrate);
+ xe_migrate_bind_queue(tile->migrate);
}
err = xe_vm_ops_add_range_unbind(&vops, range);
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (11 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 12/24] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 6:03 ` sashiko-bot
2026-09-25 4:53 ` [PATCH v7 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
` (15 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Himal Prasad Ghimiray
No reason to use the GPU for binds.
Benefits of CPU-based binds:
- Lower latency once dependencies are resolved, as there is no
interaction with the GuC or a hardware context switch both of which
are relatively slow.
- Large arrays of binds do not risk running out of migration PTEs,
avoiding -ENOBUFS being returned to userspace.
- Kernel binds are decoupled from the migration exec queue (which issues
copies and clears), so they cannot get stuck behind unrelated
jobs—this can be a problem with parallel GPU faults.
- Paves the for path decouping binds from tiles and individual engines
- Enables ULLS on the migration exec queue, as this queue has exclusive
access to the paging copy engine.
Update migration layer to formulate a PT job which will issue CPU bind
in the submission backend.
All code related to GPU-based binding has been removed.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
drivers/gpu/drm/xe/xe_bo_types.h | 2 -
drivers/gpu/drm/xe/xe_migrate.c | 249 +++----------------------------
drivers/gpu/drm/xe/xe_pt.c | 1 -
3 files changed, 17 insertions(+), 235 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_bo_types.h b/drivers/gpu/drm/xe/xe_bo_types.h
index 0eb93052c1d5..8ec4a01a0092 100644
--- a/drivers/gpu/drm/xe/xe_bo_types.h
+++ b/drivers/gpu/drm/xe/xe_bo_types.h
@@ -90,8 +90,6 @@ struct xe_bo {
/** @freed: List node for delayed put. */
struct llist_node freed;
- /** @update_index: Update index if PT BO */
- int update_index;
/** @created: Whether the bo has passed initial creation */
bool created;
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index d8d35671a190..1fabc160b7d8 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -77,18 +77,12 @@ struct xe_migrate {
* Protected by @job_mutex.
*/
struct dma_fence *fence;
- /**
- * @vm_update_sa: For integrated, used to suballocate page-tables
- * out of the pt_bo.
- */
- struct drm_suballoc_manager vm_update_sa;
/** @min_chunk_size: For dgfx, Minimum chunk size */
u64 min_chunk_size;
};
#define MAX_PREEMPTDISABLE_TRANSFER SZ_8M /* Around 1ms. */
#define MAX_CCS_LIMITED_TRANSFER SZ_4M /* XE_PAGE_SIZE * (FIELD_MAX(XE2_CCS_SIZE_MASK) + 1) */
-#define NUM_KERNEL_PDE 15
#define NUM_PT_SLOTS 48
#define LEVEL0_PAGE_TABLE_ENCODE_SIZE SZ_2M
#define MAX_NUM_PTE 512
@@ -113,7 +107,6 @@ static void xe_migrate_fini(void *arg)
dma_fence_put(m->fence);
xe_bo_put(m->pt_bo);
- drm_suballoc_manager_fini(&m->vm_update_sa);
mutex_destroy(&m->job_mutex);
xe_vm_close_and_put(m->q->vm);
xe_exec_queue_put(m->q);
@@ -234,8 +227,6 @@ static int xe_migrate_pt_bo_alloc(struct xe_tile *tile, struct xe_migrate *m,
BUILD_BUG_ON(NUM_PT_SLOTS > SZ_2M/XE_PAGE_SIZE);
/* Must be a multiple of 64K to support all platforms */
BUILD_BUG_ON(NUM_PT_SLOTS * XE_PAGE_SIZE % SZ_64K);
- /* And one slot reserved for the 4KiB page table updates */
- BUILD_BUG_ON(!(NUM_KERNEL_PDE & 1));
/* Need to be sure everything fits in the first PT, or create more */
xe_tile_assert(tile, m->batch_base_ofs + xe_bo_size(batch) < SZ_2M);
@@ -391,17 +382,9 @@ static void xe_migrate_prepare_vm(struct xe_tile *tile, struct xe_migrate *m,
}
}
- if (ofs)
- *ofs = map_ofs;
-}
-
-static void xe_migrate_suballoc_manager_init(struct xe_migrate *m, u32 map_ofs)
-{
/*
* Example layout created above, with root level = 3:
* [PT0...PT7]: kernel PT's for copy/clear; 64 or 4KiB PTE's
- * [PT8]: Kernel PT for VM_BIND, 4 KiB PTE's
- * [PT9...PT40]: Userspace PT's for VM_BIND, 4 KiB PTE's
* [PT41 = PDE 0] [PT44...PT47 = 4K and 2M vram identity maps]
*
* This makes the lowest part of the VM point to the pagetables.
@@ -409,19 +392,13 @@ static void xe_migrate_suballoc_manager_init(struct xe_migrate *m, u32 map_ofs)
* and flushes, other parts of the VM can be used either for copying and
* clearing.
*
- * For performance, the kernel reserves PDE's, so about 20 are left
- * for async VM updates.
- *
* To make it easier to work, each scratch PT is put in slot (1 + PT #)
* everywhere, this allows lockless updates to scratch pages by using
* the different addresses in VM.
*/
-#define NUM_VMUSA_UNIT_PER_PAGE 32
-#define VM_SA_UPDATE_UNIT_SIZE (XE_PAGE_SIZE / NUM_VMUSA_UNIT_PER_PAGE)
-#define NUM_VMUSA_WRITES_PER_UNIT (VM_SA_UPDATE_UNIT_SIZE / sizeof(u64))
- drm_suballoc_manager_init(&m->vm_update_sa,
- (size_t)(map_ofs / XE_PAGE_SIZE - NUM_KERNEL_PDE) *
- NUM_VMUSA_UNIT_PER_PAGE, 0);
+
+ if (ofs)
+ *ofs = map_ofs;
}
static bool xe_migrate_needs_ccs_emit(struct xe_device *xe)
@@ -466,7 +443,6 @@ static int xe_migrate_lock_prepare_vm(struct xe_tile *tile, struct xe_migrate *m
return err;
xe_migrate_prepare_vm(tile, m, vm, &map_ofs);
- xe_migrate_suballoc_manager_init(m, map_ofs);
drm_exec_retry_on_contention(&exec);
xe_validation_retry_on_oom(&ctx, &err);
}
@@ -1169,6 +1145,9 @@ struct xe_lrc *xe_migrate_lrc(struct xe_migrate *migrate)
return migrate->q->lrc[0];
}
+/* XXX: With CPU binds this can be removed in a follow up */
+#define NUM_KERNEL_PDE 15
+
static u64 migrate_vm_ppgtt_addr_tlb_inval(void)
{
/*
@@ -1788,56 +1767,6 @@ struct dma_fence *xe_migrate_clear(struct xe_migrate *m,
return fence;
}
-static void write_pgtable(struct xe_tile *tile, struct xe_bb *bb, u64 ppgtt_ofs,
- const struct xe_vm_pgtable_update_op *pt_op,
- const struct xe_vm_pgtable_update *update,
- struct xe_migrate_pt_update *pt_update)
-{
- const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
- struct xe_vm *vm = pt_update->vops->vm;
- u32 chunk;
- u32 ofs = update->ofs, size = update->qwords;
-
- /*
- * If we have 512 entries (max), we would populate it ourselves,
- * and update the PDE above it to the new pointer.
- * The only time this can only happen if we have to update the top
- * PDE. This requires a BO that is almost vm->size big.
- *
- * This shouldn't be possible in practice.. might change when 16K
- * pages are used. Hence the assert.
- */
- xe_tile_assert(tile, update->qwords < MAX_NUM_PTE);
- if (!ppgtt_ofs)
- ppgtt_ofs = xe_migrate_vram_ofs(tile_to_xe(tile),
- xe_bo_addr(update->pt_bo, 0,
- XE_PAGE_SIZE), false);
-
- do {
- u64 addr = ppgtt_ofs + ofs * 8;
-
- chunk = min(size, MAX_PTE_PER_SDI);
-
- /* Ensure populatefn can do memset64 by aligning bb->cs */
- if (!(bb->len & 1))
- bb->cs[bb->len++] = MI_NOOP;
-
- bb->cs[bb->len++] = MI_STORE_DATA_IMM | MI_SDI_NUM_QW(chunk);
- bb->cs[bb->len++] = lower_32_bits(addr);
- bb->cs[bb->len++] = upper_32_bits(addr);
- if (pt_op->bind)
- ops->populate(tile, NULL, bb->cs + bb->len,
- ofs, chunk, update);
- else
- ops->clear(vm, tile, NULL, bb->cs + bb->len,
- ofs, chunk, update);
-
- bb->len += chunk * 2;
- ofs += chunk;
- size -= chunk;
- } while (size);
-}
-
struct xe_vm *xe_migrate_get_vm(struct xe_migrate *m)
{
return xe_vm_get(m->q->vm);
@@ -1937,162 +1866,18 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
{
const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
struct xe_tile *tile = m->tile;
- struct xe_gt *gt = tile->primary_gt;
- struct xe_device *xe = tile_to_xe(tile);
struct xe_sched_job *job;
struct dma_fence *fence;
- struct drm_suballoc *sa_bo = NULL;
- struct xe_bb *bb;
- u32 i, j, batch_size = 0, ppgtt_ofs, update_idx, page_ofs = 0;
- u32 num_updates = 0, current_update = 0;
- u64 addr;
- int err = 0;
bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
- bool usm = is_migrate && xe->info.has_usm;
-
- for (i = 0; i < pt_update_ops->num_ops; ++i) {
- struct xe_vm_pgtable_update_op *pt_op = &pt_update_ops->pt_job_ops->ops[i];
- struct xe_vm_pgtable_update *updates = pt_op->entries;
-
- num_updates += pt_op->num_entries;
- for (j = 0; j < pt_op->num_entries; ++j) {
- u32 num_cmds = DIV_ROUND_UP(updates[j].qwords,
- MAX_PTE_PER_SDI);
-
- /* align noop + MI_STORE_DATA_IMM cmd prefix */
- batch_size += 4 * num_cmds + updates[j].qwords * 2;
- }
- }
-
- /* fixed + PTE entries */
- if (IS_DGFX(xe))
- batch_size += 2;
- else
- batch_size += 6 * (num_updates / MAX_PTE_PER_SDI + 1) +
- num_updates * 2;
-
- bb = xe_bb_new(gt, batch_size, usm);
- if (IS_ERR(bb))
- return ERR_CAST(bb);
-
- /* For sysmem PTE's, need to map them in our hole.. */
- if (!IS_DGFX(xe)) {
- u16 pat_index = xe_cache_pat_idx(xe, XE_CACHE_WB);
- u32 ptes, ofs;
-
- ppgtt_ofs = NUM_KERNEL_PDE - 1;
- if (!is_migrate) {
- u32 num_units = DIV_ROUND_UP(num_updates,
- NUM_VMUSA_WRITES_PER_UNIT);
-
- if (num_units > m->vm_update_sa.size) {
- err = -ENOBUFS;
- goto err_bb;
- }
- sa_bo = drm_suballoc_new(&m->vm_update_sa, num_units,
- GFP_KERNEL, true, 0);
- if (IS_ERR(sa_bo)) {
- err = PTR_ERR(sa_bo);
- goto err_bb;
- }
-
- ppgtt_ofs = NUM_KERNEL_PDE +
- (drm_suballoc_soffset(sa_bo) /
- NUM_VMUSA_UNIT_PER_PAGE);
- page_ofs = (drm_suballoc_soffset(sa_bo) %
- NUM_VMUSA_UNIT_PER_PAGE) *
- VM_SA_UPDATE_UNIT_SIZE;
- }
-
- /* Map our PT's to gtt */
- i = 0;
- j = 0;
- ptes = num_updates;
- ofs = ppgtt_ofs * XE_PAGE_SIZE + page_ofs;
- while (ptes) {
- u32 chunk = min(MAX_PTE_PER_SDI, ptes);
- u32 idx = 0;
-
- bb->cs[bb->len++] = MI_STORE_DATA_IMM |
- MI_SDI_NUM_QW(chunk);
- bb->cs[bb->len++] = ofs;
- bb->cs[bb->len++] = 0; /* upper_32_bits */
-
- for (; i < pt_update_ops->num_ops; ++i) {
- struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->pt_job_ops->ops[i];
- struct xe_vm_pgtable_update *updates = pt_op->entries;
-
- for (; j < pt_op->num_entries; ++j, ++current_update, ++idx) {
- struct xe_vm *vm = pt_update->vops->vm;
- struct xe_bo *pt_bo = updates[j].pt_bo;
-
- if (idx == chunk)
- goto next_cmd;
-
- xe_tile_assert(tile, xe_bo_size(pt_bo) == SZ_4K);
-
- /* Map a PT at most once */
- if (pt_bo->update_index < 0)
- pt_bo->update_index = current_update;
-
- addr = vm->pt_ops->pte_encode_bo(pt_bo, 0,
- pat_index, 0);
- bb->cs[bb->len++] = lower_32_bits(addr);
- bb->cs[bb->len++] = upper_32_bits(addr);
- }
-
- j = 0;
- }
-
-next_cmd:
- ptes -= chunk;
- ofs += chunk * sizeof(u64);
- }
-
- bb->cs[bb->len++] = MI_BATCH_BUFFER_END;
- update_idx = bb->len;
-
- addr = xe_migrate_vm_addr(ppgtt_ofs, 0) +
- (page_ofs / sizeof(u64)) * XE_PAGE_SIZE;
- for (i = 0; i < pt_update_ops->num_ops; ++i) {
- struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->pt_job_ops->ops[i];
- struct xe_vm_pgtable_update *updates = pt_op->entries;
-
- for (j = 0; j < pt_op->num_entries; ++j) {
- struct xe_bo *pt_bo = updates[j].pt_bo;
-
- write_pgtable(tile, bb, addr +
- pt_bo->update_index * XE_PAGE_SIZE,
- pt_op, &updates[j], pt_update);
- }
- }
- } else {
- /* phys pages, no preamble required */
- bb->cs[bb->len++] = MI_BATCH_BUFFER_END;
- update_idx = bb->len;
-
- for (i = 0; i < pt_update_ops->num_ops; ++i) {
- struct xe_vm_pgtable_update_op *pt_op =
- &pt_update_ops->pt_job_ops->ops[i];
- struct xe_vm_pgtable_update *updates = pt_op->entries;
-
- for (j = 0; j < pt_op->num_entries; ++j)
- write_pgtable(tile, bb, 0, pt_op, &updates[j],
- pt_update);
- }
- }
+ int err;
- job = xe_bb_create_migration_job(pt_update_ops->q, bb,
- xe_migrate_batch_base(m, usm),
- update_idx);
+ job = xe_sched_job_create(pt_update_ops->q, NULL);
if (IS_ERR(job)) {
err = PTR_ERR(job);
- goto err_sa;
+ goto err_out;
}
- xe_sched_job_add_migrate_flush(job, MI_INVALIDATE_TLB);
+ xe_tile_assert(tile, job->is_pt_job);
if (ops->pre_commit) {
pt_update->job = job;
@@ -2103,6 +1888,12 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
if (is_migrate)
mutex_lock(&m->job_mutex);
+ job->pt_update[0].vm = pt_update->vops->vm;
+ job->pt_update[0].tile = tile;
+ job->pt_update[0].ops = ops;
+ job->pt_update[0].pt_job_ops =
+ xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
+
xe_sched_job_arm(job);
fence = dma_fence_get(&job->drm.s_fence->finished);
xe_sched_job_push(job);
@@ -2110,17 +1901,11 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
if (is_migrate)
mutex_unlock(&m->job_mutex);
- xe_bb_free(bb, fence);
- drm_suballoc_free(sa_bo, fence);
-
return fence;
err_job:
xe_sched_job_put(job);
-err_sa:
- drm_suballoc_free(sa_bo, NULL);
-err_bb:
- xe_bb_free(bb, NULL);
+err_out:
return ERR_PTR(err);
}
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index a483158f94d4..666097d68563 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -379,7 +379,6 @@ xe_pt_new_shared(struct xe_walk_update *wupd, struct xe_pt *parent,
entry->pt = parent;
entry->flags = 0;
entry->qwords = 0;
- entry->pt_bo->update_index = -1;
entry->level = parent->level;
if (alloc_entries) {
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs
2026-09-25 4:53 ` [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
@ 2026-09-25 6:03 ` sashiko-bot
2026-09-25 6:54 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 6:03 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] drm/xe: memory leak of pt_job_ops and dma_fence when PT jobs are canceled
- [High] drm/xe: GPU hang and memory leak due to missing CPU bind support in Execlist backend
--
commit 6471c15e2149c2517458e891e7dceef8c5573f97
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: Enable CPU binds for jobs
This commit updates the migration layer to formulate page table jobs that
issue CPU binds in the submission backend. All code related to GPU-based
binding is removed to lower latency and avoid PTE migration limits.
> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
[ ... ]
> @@ -1937,162 +1866,18 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
> {
> const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
> struct xe_tile *tile = m->tile;
> struct xe_sched_job *job;
> struct dma_fence *fence;
> bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
> + int err;
>
> - job = xe_bb_create_migration_job(pt_update_ops->q, bb,
> - xe_migrate_batch_base(m, usm),
> - update_idx);
> + job = xe_sched_job_create(pt_update_ops->q, NULL);
[Severity: High]
Does the Execlist backend properly handle these new PT jobs?
Since xe_sched_job_create() is now called with a NULL batch address, these
jobs rely on backend interception. However, looking at execlist_run_job() in
the Execlist backend:
drivers/gpu/drm/xe/xe_execlist.c:execlist_run_job() {
...
q->ring_ops->emit_job(job);
xe_execlist_make_active(exl);
return job->fence;
}
If GuC submission is disabled, it appears the backend receives this job and
directly emits it to the ring without an is_pt_job() check. Could this result
in pushing an uninitialized batch address to the hardware, causing a GPU hang
and leaking the pt_job_ops structures?
> if (IS_ERR(job)) {
> err = PTR_ERR(job);
> - goto err_sa;
> + goto err_out;
> }
>
> - xe_sched_job_add_migrate_flush(job, MI_INVALIDATE_TLB);
> + xe_tile_assert(tile, job->is_pt_job);
>
> if (ops->pre_commit) {
> pt_update->job = job;
> @@ -2103,6 +1888,12 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
> if (is_migrate)
> mutex_lock(&m->job_mutex);
>
> + job->pt_update[0].vm = pt_update->vops->vm;
> + job->pt_update[0].tile = tile;
> + job->pt_update[0].ops = ops;
> + job->pt_update[0].pt_job_ops =
> + xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
> +
[Severity: High]
Is there a memory leak if this PT job gets canceled before execution?
We acquire a reference to pt_job_ops here, and xe_sched_job_arm() takes an
extra dma_fence reference. If the job is canceled (for example, from an
execution queue kill or GPU reset), the DRM scheduler calls the free_job
callback which ultimately calls xe_sched_job_destroy():
drivers/gpu/drm/xe/xe_sched_job.c:xe_sched_job_destroy() {
...
xe_sched_job_free_fences(job);
dma_fence_put(job->fence);
drm_sched_job_cleanup(&job->drm);
job_free(job);
}
Since xe_sched_job_destroy() doesn't check job->is_pt_job, it seems we never
drop the pt_job_ops or the extra dma_fence reference. This would permanently
leak the deferred actions and their associated page table BO references. Does
the job destruction path need updating to clean up canceled PT jobs?
> xe_sched_job_arm(job);
> fence = dma_fence_get(&job->drm.s_fence->finished);
> xe_sched_job_push(job);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=13
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs
2026-09-25 6:03 ` sashiko-bot
@ 2026-09-25 6:54 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 6:54 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 06:03:25AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [High] drm/xe: memory leak of pt_job_ops and dma_fence when PT jobs are canceled
> - [High] drm/xe: GPU hang and memory leak due to missing CPU bind support in Execlist backend
> --
>
> commit 6471c15e2149c2517458e891e7dceef8c5573f97
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Enable CPU binds for jobs
>
> This commit updates the migration layer to formulate page table jobs that
> issue CPU binds in the submission backend. All code related to GPU-based
> binding is removed to lower latency and avoid PTE migration limits.
>
> > diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> > --- a/drivers/gpu/drm/xe/xe_migrate.c
> > +++ b/drivers/gpu/drm/xe/xe_migrate.c
> [ ... ]
> > @@ -1937,162 +1866,18 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
> > {
> > const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
> > struct xe_tile *tile = m->tile;
> > struct xe_sched_job *job;
> > struct dma_fence *fence;
> > bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
> > + int err;
> >
> > - job = xe_bb_create_migration_job(pt_update_ops->q, bb,
> > - xe_migrate_batch_base(m, usm),
> > - update_idx);
> > + job = xe_sched_job_create(pt_update_ops->q, NULL);
>
> [Severity: High]
> Does the Execlist backend properly handle these new PT jobs?
>
Execlist backend is dead. Non-issue.
> Since xe_sched_job_create() is now called with a NULL batch address, these
> jobs rely on backend interception. However, looking at execlist_run_job() in
> the Execlist backend:
>
> drivers/gpu/drm/xe/xe_execlist.c:execlist_run_job() {
> ...
> q->ring_ops->emit_job(job);
> xe_execlist_make_active(exl);
>
> return job->fence;
> }
>
> If GuC submission is disabled, it appears the backend receives this job and
> directly emits it to the ring without an is_pt_job() check. Could this result
> in pushing an uninitialized batch address to the hardware, causing a GPU hang
> and leaking the pt_job_ops structures?
>
> > if (IS_ERR(job)) {
> > err = PTR_ERR(job);
> > - goto err_sa;
> > + goto err_out;
> > }
> >
> > - xe_sched_job_add_migrate_flush(job, MI_INVALIDATE_TLB);
> > + xe_tile_assert(tile, job->is_pt_job);
> >
> > if (ops->pre_commit) {
> > pt_update->job = job;
> > @@ -2103,6 +1888,12 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
> > if (is_migrate)
> > mutex_lock(&m->job_mutex);
> >
> > + job->pt_update[0].vm = pt_update->vops->vm;
> > + job->pt_update[0].tile = tile;
> > + job->pt_update[0].ops = ops;
> > + job->pt_update[0].pt_job_ops =
> > + xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
> > +
>
> [Severity: High]
> Is there a memory leak if this PT job gets canceled before execution?
>
Again non-issue - _arm() always equals run_job() call. Non-issue.
Matt
> We acquire a reference to pt_job_ops here, and xe_sched_job_arm() takes an
> extra dma_fence reference. If the job is canceled (for example, from an
> execution queue kill or GPU reset), the DRM scheduler calls the free_job
> callback which ultimately calls xe_sched_job_destroy():
>
> drivers/gpu/drm/xe/xe_sched_job.c:xe_sched_job_destroy() {
> ...
> xe_sched_job_free_fences(job);
> dma_fence_put(job->fence);
> drm_sched_job_cleanup(&job->drm);
> job_free(job);
> }
>
> Since xe_sched_job_destroy() doesn't check job->is_pt_job, it seems we never
> drop the pt_job_ops or the extra dma_fence reference. This would permanently
> leak the deferred actions and their associated page table BO references. Does
> the job destruction path need updating to clean up canceled PT jobs?
>
> > xe_sched_job_arm(job);
> > fence = dma_fence_get(&job->drm.s_fence->finished);
> > xe_sched_job_push(job);
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=13
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (12 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 15/24] drm/xe: Make bind queues operate cross-tile Matthew Brost
` (14 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Shuicheng Lin
Both populate and clear have unused void* ptr arguments, remove these.
Also consolidate ofs, num_qwords in update argument.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
---
v7:
- Drop ofs, num_qwords (Shuicheng)
---
drivers/gpu/drm/xe/xe_migrate.c | 2 --
drivers/gpu/drm/xe/xe_migrate.h | 10 +-------
drivers/gpu/drm/xe/xe_pt.c | 41 +++++++++++----------------------
3 files changed, 15 insertions(+), 38 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 1fabc160b7d8..1c16d31f0d10 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -1811,11 +1811,9 @@ xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
if (pt_op->bind && !force_clear)
ops->populate(tile, &update->pt_bo->vmap,
- NULL, update->ofs, update->qwords,
update);
else
ops->clear(vm, tile, &update->pt_bo->vmap,
- NULL, update->ofs, update->qwords,
update);
}
}
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index e6b7070134e1..94b8d8d63720 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -43,9 +43,6 @@ struct xe_migrate_pt_update_ops {
* @populate: Populate a command buffer or page-table with ptes.
* @tile: The tile for the current operation.
* @map: struct iosys_map into the memory to be populated.
- * @pos: If @map is NULL, map into the memory to be populated.
- * @ofs: qword offset into @map, unused if @map is NULL.
- * @num_qwords: Number of qwords to write.
* @update: Information about the PTEs to be inserted.
*
* This interface is intended to be used as a callback into the
@@ -53,16 +50,12 @@ struct xe_migrate_pt_update_ops {
* page-tables with PTEs.
*/
void (*populate)(struct xe_tile *tile, struct iosys_map *map,
- void *pos, u32 ofs, u32 num_qwords,
const struct xe_vm_pgtable_update *update);
/**
* @clear: Clear a command buffer or page-table with ptes.
* @vm: VM being updated
* @tile: The tile for the current operation.
* @map: struct iosys_map into the memory to be populated.
- * @pos: If @map is NULL, map into the memory to be populated.
- * @ofs: qword offset into @map, unused if @map is NULL.
- * @num_qwords: Number of qwords to write.
* @update: Information about the PTEs to be inserted.
*
* This interface is intended to be used as a callback into the
@@ -70,8 +63,7 @@ struct xe_migrate_pt_update_ops {
* page-tables with PTEs.
*/
void (*clear)(struct xe_vm *vm, struct xe_tile *tile,
- struct iosys_map *map, void *pos, u32 ofs,
- u32 num_qwords,
+ struct iosys_map *map,
const struct xe_vm_pgtable_update *update);
/**
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index 666097d68563..102f40c966f1 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -1106,30 +1106,17 @@ bool xe_pt_zap_ptes_range(struct xe_tile *tile, struct xe_vm *vm,
static void
xe_vm_populate_pgtable(struct xe_tile *tile, struct iosys_map *map,
- void *data, u32 qword_ofs, u32 num_qwords,
const struct xe_vm_pgtable_update *update)
{
struct xe_pt_entry *ptes = update->pt_entries;
- u64 *ptr = data;
u32 i;
- /*
- * @qword_ofs is the absolute entry offset within the page table, while
- * @ptes is indexed relative to @update->ofs (its first entry). The GPU
- * path (write_pgtable) splits a single update into MAX_PTE_PER_SDI-sized
- * chunks, calling this with an advancing @qword_ofs but a fresh @data
- * pointer per chunk, so translate back into a @ptes index rather than
- * assuming the chunk starts at ptes[0].
- */
- for (i = 0; i < num_qwords; i++) {
- u32 idx = qword_ofs - update->ofs + i;
+ xe_assert(tile_to_xe(tile), map);
+ xe_assert(tile_to_xe(tile), !iosys_map_is_null(map));
- if (map)
- xe_map_wr(tile_to_xe(tile), map, (qword_ofs + i) *
- sizeof(u64), u64, ptes[idx].pte);
- else
- ptr[i] = ptes[idx].pte;
- }
+ for (i = 0; i < update->qwords; i++)
+ xe_map_wr(tile_to_xe(tile), map, (update->ofs + i) *
+ sizeof(u64), u64, ptes[i].pte);
}
static void xe_pt_cancel_bind(struct xe_vma *vma,
@@ -2021,22 +2008,22 @@ static unsigned int xe_pt_stage_unbind(struct xe_tile *tile,
static void
xe_migrate_clear_pgtable_callback(struct xe_vm *vm, struct xe_tile *tile,
- struct iosys_map *map, void *ptr,
- u32 qword_ofs, u32 num_qwords,
+ struct iosys_map *map,
const struct xe_vm_pgtable_update *update)
{
u64 empty = __xe_pt_empty_pte(tile, vm, update->level);
int i;
- if (map && map->is_iomem)
- for (i = 0; i < num_qwords; ++i)
- xe_map_wr(tile_to_xe(tile), map, (qword_ofs + i) *
+ xe_assert(vm->xe, map);
+ xe_assert(vm->xe, !iosys_map_is_null(map));
+
+ if (map->is_iomem)
+ for (i = 0; i < update->qwords; ++i)
+ xe_map_wr(tile_to_xe(tile), map, (update->ofs + i) *
sizeof(u64), u64, empty);
- else if (map)
- memset64(map->vaddr + qword_ofs * sizeof(u64), empty,
- num_qwords);
else
- memset64(ptr, empty, num_qwords);
+ memset64(map->vaddr + update->ofs * sizeof(u64), empty,
+ update->qwords);
}
static void xe_pt_abort_unbind(struct xe_vma *vma,
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 15/24] drm/xe: Make bind queues operate cross-tile
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (13 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 16/24] drm/xe: Add CPU bind layer Matthew Brost
` (13 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Shuicheng Lin
Since bind jobs execute on the CPU rather than the GPU, maintaining a
per-tile bind queue no longer provides value. Convert the driver to use
a single bind queue shared across tiles. The primary change is routing
all GT TLB invalidations through this unified bind queue.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
---
v7:
- to_dep_scheduler drop tile argument (Shuicheng)
- gt_to_xe vs q->vm->xe (Shuicheng)
---
drivers/gpu/drm/xe/xe_exec_queue.c | 142 +++++++++--------------
drivers/gpu/drm/xe/xe_exec_queue.h | 14 +--
drivers/gpu/drm/xe/xe_exec_queue_types.h | 20 +---
drivers/gpu/drm/xe/xe_pt.c | 21 ++--
drivers/gpu/drm/xe/xe_sync.c | 20 +---
drivers/gpu/drm/xe/xe_tlb_inval_job.c | 15 ++-
drivers/gpu/drm/xe/xe_tlb_inval_job.h | 2 +-
drivers/gpu/drm/xe/xe_vm.c | 65 +++++------
drivers/gpu/drm/xe/xe_vm_types.h | 2 +-
9 files changed, 124 insertions(+), 177 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.c b/drivers/gpu/drm/xe/xe_exec_queue.c
index 7bbb31431900..d2270f32e35d 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.c
+++ b/drivers/gpu/drm/xe/xe_exec_queue.c
@@ -142,9 +142,8 @@ static void __xe_exec_queue_free(struct xe_exec_queue *q)
{
int i;
- for (i = 0; i < XE_EXEC_QUEUE_TLB_INVAL_COUNT; ++i)
- if (q->tlb_inval[i].dep_scheduler)
- xe_dep_scheduler_fini(q->tlb_inval[i].dep_scheduler);
+ for_each_tlb_inval(q, i)
+ xe_dep_scheduler_fini(q->tlb_inval[i].dep_scheduler);
if (xe_exec_queue_uses_pxp(q))
xe_pxp_exec_queue_remove(gt_to_xe(q->gt)->pxp, q);
@@ -166,31 +165,34 @@ static void __xe_exec_queue_free(struct xe_exec_queue *q)
static int alloc_dep_schedulers(struct xe_device *xe, struct xe_exec_queue *q)
{
- struct xe_tile *tile = gt_to_tile(q->gt);
- int i;
+ struct xe_tile *tile;
+ int i = 0, j;
+ u8 id;
- for (i = 0; i < XE_EXEC_QUEUE_TLB_INVAL_COUNT; ++i) {
- struct xe_dep_scheduler *dep_scheduler;
- struct xe_gt *gt;
- struct workqueue_struct *wq;
+ for_each_tile(tile, xe, id) {
+ for (j = 0; j < (XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT + 1); ++j, ++i) {
+ struct xe_dep_scheduler *dep_scheduler;
+ struct xe_gt *gt;
+ struct workqueue_struct *wq;
- if (i == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT)
- gt = tile->primary_gt;
- else
- gt = tile->media_gt;
+ if (j == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT)
+ gt = tile->primary_gt;
+ else
+ gt = tile->media_gt;
- if (!gt)
- continue;
+ if (!gt)
+ continue;
- wq = gt->tlb_inval.job_wq;
+ wq = gt->tlb_inval.job_wq;
#define MAX_TLB_INVAL_JOBS 16 /* Picking a reasonable value */
- dep_scheduler = xe_dep_scheduler_create(xe, wq, q->name,
- MAX_TLB_INVAL_JOBS);
- if (IS_ERR(dep_scheduler))
- return PTR_ERR(dep_scheduler);
+ dep_scheduler = xe_dep_scheduler_create(xe, wq, q->name,
+ MAX_TLB_INVAL_JOBS);
+ if (IS_ERR(dep_scheduler))
+ return PTR_ERR(dep_scheduler);
- q->tlb_inval[i].dep_scheduler = dep_scheduler;
+ q->tlb_inval[i].dep_scheduler = dep_scheduler;
+ }
}
#undef MAX_TLB_INVAL_JOBS
@@ -224,7 +226,6 @@ static struct xe_exec_queue *__xe_exec_queue_alloc(struct xe_device *xe,
q->ops = gt->exec_queue_ops;
INIT_LIST_HEAD(&q->lr.link);
INIT_LIST_HEAD(&q->vm_exec_queue_link);
- INIT_LIST_HEAD(&q->multi_gt_link);
INIT_LIST_HEAD(&q->hw_engine_group_link);
INIT_LIST_HEAD(&q->pxp.link);
spin_lock_init(&q->multi_queue.lock);
@@ -636,7 +637,6 @@ ALLOW_ERROR_INJECTION(xe_exec_queue_create_bind, ERRNO);
void xe_exec_queue_destroy(struct kref *ref)
{
struct xe_exec_queue *q = container_of(ref, struct xe_exec_queue, refcount);
- struct xe_exec_queue *eq, *next;
int i;
xe_assert(gt_to_xe(q->gt), atomic_read(&q->job_cnt) == 0);
@@ -648,15 +648,9 @@ void xe_exec_queue_destroy(struct kref *ref)
xe_pxp_exec_queue_remove(gt_to_xe(q->gt)->pxp, q);
xe_exec_queue_last_fence_put_unlocked(q);
- for_each_tlb_inval(i)
+ for_each_tlb_inval(q, i)
xe_exec_queue_tlb_inval_last_fence_put_unlocked(q, i);
- if (!(q->flags & EXEC_QUEUE_FLAG_BIND_ENGINE_CHILD)) {
- list_for_each_entry_safe(eq, next, &q->multi_gt_list,
- multi_gt_link)
- xe_exec_queue_put(eq);
- }
-
if (q->user_vm) {
xe_vm_put(q->user_vm);
q->user_vm = NULL;
@@ -1343,7 +1337,6 @@ int xe_exec_queue_create_ioctl(struct drm_device *dev, void *data,
u64_to_user_ptr(args->instances);
struct xe_hw_engine *hwe;
struct xe_vm *vm;
- struct xe_tile *tile;
struct xe_exec_queue *q = NULL;
u32 logical_mask;
u32 flags = 0;
@@ -1392,31 +1385,16 @@ int xe_exec_queue_create_ioctl(struct drm_device *dev, void *data,
return -ENOENT;
}
- for_each_tile(tile, xe, id) {
- struct xe_exec_queue *new;
-
- flags |= EXEC_QUEUE_FLAG_VM;
- if (id)
- flags |= EXEC_QUEUE_FLAG_BIND_ENGINE_CHILD;
-
- new = xe_exec_queue_create_bind(xe, tile, vm, flags,
- args->extensions);
- if (IS_ERR(new)) {
- up_read(&vm->lock);
- xe_vm_put(vm);
- err = PTR_ERR(new);
- if (q)
- goto put_exec_queue;
- return err;
- }
- if (id == 0)
- q = new;
- else
- list_add_tail(&new->multi_gt_list,
- &q->multi_gt_link);
- }
+ flags |= EXEC_QUEUE_FLAG_VM;
+
+ q = xe_exec_queue_create_bind(xe, xe_device_get_root_tile(xe),
+ vm, flags, args->extensions);
up_read(&vm->lock);
xe_vm_put(vm);
+ if (IS_ERR(q)) {
+ err = PTR_ERR(q);
+ return err;
+ }
} else {
logical_mask = calc_validate_logical_mask(xe, eci,
args->width,
@@ -1638,14 +1616,6 @@ void xe_exec_queue_update_run_ticks(struct xe_exec_queue *q)
*/
void xe_exec_queue_kill(struct xe_exec_queue *q)
{
- struct xe_exec_queue *eq = q, *next;
-
- list_for_each_entry_safe(eq, next, &eq->multi_gt_list,
- multi_gt_link) {
- q->ops->kill(eq);
- xe_vm_remove_compute_exec_queue(q->vm, eq);
- }
-
q->ops->kill(q);
xe_vm_remove_compute_exec_queue(q->vm, q);
}
@@ -1806,42 +1776,40 @@ void xe_exec_queue_last_fence_set(struct xe_exec_queue *q, struct xe_vm *vm,
* xe_exec_queue_tlb_inval_last_fence_put() - Drop ref to last TLB invalidation fence
* @q: The exec queue
* @vm: The VM the engine does a bind for
- * @type: Either primary or media GT
+ * @idx: Index of tlb invalidation
*/
void xe_exec_queue_tlb_inval_last_fence_put(struct xe_exec_queue *q,
struct xe_vm *vm,
- unsigned int type)
+ unsigned int idx)
{
xe_exec_queue_last_fence_lockdep_assert(q, vm);
- xe_assert(vm->xe, type == XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT ||
- type == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT);
+ xe_assert(vm->xe, idx < XE_EXEC_QUEUE_TLB_INVAL_COUNT);
- xe_exec_queue_tlb_inval_last_fence_put_unlocked(q, type);
+ xe_exec_queue_tlb_inval_last_fence_put_unlocked(q, idx);
}
/**
* xe_exec_queue_tlb_inval_last_fence_put_unlocked() - Drop ref to last TLB
* invalidation fence unlocked
* @q: The exec queue
- * @type: Either primary or media GT
+ * @idx: Index of tlb invalidation
*
* Only safe to be called from xe_exec_queue_destroy().
*/
void xe_exec_queue_tlb_inval_last_fence_put_unlocked(struct xe_exec_queue *q,
- unsigned int type)
+ unsigned int idx)
{
- xe_assert(gt_to_xe(q->gt), type == XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT ||
- type == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT);
+ xe_assert(gt_to_xe(q->gt), idx < XE_EXEC_QUEUE_TLB_INVAL_COUNT);
- dma_fence_put(q->tlb_inval[type].last_fence);
- q->tlb_inval[type].last_fence = NULL;
+ dma_fence_put(q->tlb_inval[idx].last_fence);
+ q->tlb_inval[idx].last_fence = NULL;
}
/**
* xe_exec_queue_tlb_inval_last_fence_get() - Get last fence for TLB invalidation
* @q: The exec queue
* @vm: The VM the engine does a bind for
- * @type: Either primary or media GT
+ * @idx: Index of tlb invalidation
*
* Get last fence, takes a ref
*
@@ -1849,22 +1817,21 @@ void xe_exec_queue_tlb_inval_last_fence_put_unlocked(struct xe_exec_queue *q,
*/
struct dma_fence *xe_exec_queue_tlb_inval_last_fence_get(struct xe_exec_queue *q,
struct xe_vm *vm,
- unsigned int type)
+ unsigned int idx)
{
struct dma_fence *fence;
xe_exec_queue_last_fence_lockdep_assert(q, vm);
- xe_assert(vm->xe, type == XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT ||
- type == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT);
+ xe_assert(vm->xe, idx < XE_EXEC_QUEUE_TLB_INVAL_COUNT);
xe_assert(vm->xe, q->flags & (EXEC_QUEUE_FLAG_VM |
EXEC_QUEUE_FLAG_MIGRATE));
- if (q->tlb_inval[type].last_fence &&
+ if (q->tlb_inval[idx].last_fence &&
test_bit(DMA_FENCE_FLAG_SIGNALED_BIT,
- &q->tlb_inval[type].last_fence->flags))
- xe_exec_queue_tlb_inval_last_fence_put(q, vm, type);
+ &q->tlb_inval[idx].last_fence->flags))
+ xe_exec_queue_tlb_inval_last_fence_put(q, vm, idx);
- fence = q->tlb_inval[type].last_fence ?: dma_fence_get_stub();
+ fence = q->tlb_inval[idx].last_fence ?: dma_fence_get_stub();
dma_fence_get(fence);
return fence;
}
@@ -1874,26 +1841,25 @@ struct dma_fence *xe_exec_queue_tlb_inval_last_fence_get(struct xe_exec_queue *q
* @q: The exec queue
* @vm: The VM the engine does a bind for
* @fence: The fence
- * @type: Either primary or media GT
+ * @idx: Index of tlb invalidation
*
- * Set the last fence for the tlb invalidation type on the queue. Increases
+ * Set the last fence for the tlb invalidation client on the queue. Increases
* reference count for fence, when closing queue
* xe_exec_queue_tlb_inval_last_fence_put should be called.
*/
void xe_exec_queue_tlb_inval_last_fence_set(struct xe_exec_queue *q,
struct xe_vm *vm,
struct dma_fence *fence,
- unsigned int type)
+ unsigned int idx)
{
xe_exec_queue_last_fence_lockdep_assert(q, vm);
- xe_assert(vm->xe, type == XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT ||
- type == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT);
+ xe_assert(vm->xe, idx < XE_EXEC_QUEUE_TLB_INVAL_COUNT);
xe_assert(vm->xe, q->flags & (EXEC_QUEUE_FLAG_VM |
EXEC_QUEUE_FLAG_MIGRATE));
xe_assert(vm->xe, !dma_fence_is_container(fence));
- xe_exec_queue_tlb_inval_last_fence_put(q, vm, type);
- q->tlb_inval[type].last_fence = dma_fence_get(fence);
+ xe_exec_queue_tlb_inval_last_fence_put(q, vm, idx);
+ q->tlb_inval[idx].last_fence = dma_fence_get(fence);
}
/**
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.h b/drivers/gpu/drm/xe/xe_exec_queue.h
index 0225426c57b0..b02a390ba989 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.h
+++ b/drivers/gpu/drm/xe/xe_exec_queue.h
@@ -14,9 +14,9 @@ struct drm_file;
struct xe_device;
struct xe_file;
-#define for_each_tlb_inval(__i) \
- for (__i = XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT; \
- __i <= XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT; ++__i)
+#define for_each_tlb_inval(__q, __i) \
+ for (__i = 0; __i < XE_EXEC_QUEUE_TLB_INVAL_COUNT; ++__i) \
+ for_each_if((__q)->tlb_inval[__i].dep_scheduler)
struct xe_exec_queue *xe_exec_queue_create(struct xe_device *xe, struct xe_vm *vm,
u32 logical_mask, u16 width,
@@ -141,19 +141,19 @@ void xe_exec_queue_last_fence_set(struct xe_exec_queue *e, struct xe_vm *vm,
void xe_exec_queue_tlb_inval_last_fence_put(struct xe_exec_queue *q,
struct xe_vm *vm,
- unsigned int type);
+ unsigned int idx);
void xe_exec_queue_tlb_inval_last_fence_put_unlocked(struct xe_exec_queue *q,
- unsigned int type);
+ unsigned int idx);
struct dma_fence *xe_exec_queue_tlb_inval_last_fence_get(struct xe_exec_queue *q,
struct xe_vm *vm,
- unsigned int type);
+ unsigned int idx);
void xe_exec_queue_tlb_inval_last_fence_set(struct xe_exec_queue *q,
struct xe_vm *vm,
struct dma_fence *fence,
- unsigned int type);
+ unsigned int idx);
void xe_exec_queue_update_run_ticks(struct xe_exec_queue *q);
diff --git a/drivers/gpu/drm/xe/xe_exec_queue_types.h b/drivers/gpu/drm/xe/xe_exec_queue_types.h
index 836f88fc0faa..48c05398b016 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue_types.h
+++ b/drivers/gpu/drm/xe/xe_exec_queue_types.h
@@ -137,16 +137,14 @@ struct xe_exec_queue {
#define EXEC_QUEUE_FLAG_KERNEL BIT(0)
/* for VM jobs. Caller needs to hold rpm ref when creating queue with this flag */
#define EXEC_QUEUE_FLAG_VM BIT(1)
-/* child of VM queue for multi-tile VM jobs */
-#define EXEC_QUEUE_FLAG_BIND_ENGINE_CHILD BIT(2)
/* kernel exec_queue only, set priority to highest level */
-#define EXEC_QUEUE_FLAG_HIGH_PRIORITY BIT(3)
+#define EXEC_QUEUE_FLAG_HIGH_PRIORITY BIT(2)
/* flag to indicate low latency hint to guc */
-#define EXEC_QUEUE_FLAG_LOW_LATENCY BIT(4)
+#define EXEC_QUEUE_FLAG_LOW_LATENCY BIT(3)
/* for migration (kernel copy, clear, bind) jobs */
-#define EXEC_QUEUE_FLAG_MIGRATE BIT(5)
+#define EXEC_QUEUE_FLAG_MIGRATE BIT(4)
/* for programming COMMON_SLICE_CHICKEN3 on first submission */
-#define EXEC_QUEUE_FLAG_DISABLE_STATE_CACHE_PERF_FIX BIT(6)
+#define EXEC_QUEUE_FLAG_DISABLE_STATE_CACHE_PERF_FIX BIT(5)
/**
* @flags: flags for this exec queue, should statically setup aside from ban
@@ -157,13 +155,6 @@ struct xe_exec_queue {
/** @ban_reason: Bitmask of ban reasons (DRM_XE_EXEC_QUEUE_BAN_REASON_*) */
atomic_t ban_reason;
- union {
- /** @multi_gt_list: list head for VM bind engines if multi-GT */
- struct list_head multi_gt_list;
- /** @multi_gt_link: link for VM bind engines if multi-GT */
- struct list_head multi_gt_link;
- };
-
union {
/** @execlist: execlist backend specific state for exec queue */
struct xe_execlist_exec_queue *execlist;
@@ -230,7 +221,8 @@ struct xe_exec_queue {
#define XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT 0
#define XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT 1
-#define XE_EXEC_QUEUE_TLB_INVAL_COUNT (XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT + 1)
+#define XE_EXEC_QUEUE_TLB_INVAL_COUNT \
+ ((XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT + 1) * 2)
/** @tlb_inval: TLB invalidations exec queue state */
struct {
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index 102f40c966f1..bf0168d071c0 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -2708,12 +2708,17 @@ static const struct xe_migrate_pt_update_ops svm_userptr_migrate_ops;
#endif
static struct xe_dep_scheduler *to_dep_scheduler(struct xe_exec_queue *q,
- struct xe_gt *gt)
+ struct xe_gt *gt,
+ unsigned int *type)
{
+ int tile_ofs = gt_to_tile(gt)->id * (XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT + 1);
+
if (xe_gt_is_media_type(gt))
- return q->tlb_inval[XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT].dep_scheduler;
+ *type = tile_ofs + XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT;
+ else
+ *type = tile_ofs + XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT;
- return q->tlb_inval[XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT].dep_scheduler;
+ return q->tlb_inval[*type].dep_scheduler;
}
/**
@@ -2738,6 +2743,7 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
struct xe_tlb_inval_job *ijob = NULL, *mjob = NULL;
struct xe_range_fence *rfence;
struct xe_vma_op *op;
+ unsigned int type;
int err = 0, i;
struct xe_migrate_pt_update update = {
.ops = pt_update_ops->needs_svm_lock ?
@@ -2764,13 +2770,13 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
if (pt_update_ops->needs_invalidation) {
struct xe_dep_scheduler *dep_scheduler =
- to_dep_scheduler(q, tile->primary_gt);
+ to_dep_scheduler(q, tile->primary_gt, &type);
ijob = xe_tlb_inval_job_create(q, &tile->primary_gt->tlb_inval,
dep_scheduler, vm,
pt_update_ops->start,
pt_update_ops->last,
- XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT);
+ type);
if (IS_ERR(ijob)) {
err = PTR_ERR(ijob);
goto kill_vm_tile1;
@@ -2789,14 +2795,15 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
}
if (tile->media_gt) {
- dep_scheduler = to_dep_scheduler(q, tile->media_gt);
+ dep_scheduler = to_dep_scheduler(q, tile->media_gt,
+ &type);
mjob = xe_tlb_inval_job_create(q,
&tile->media_gt->tlb_inval,
dep_scheduler, vm,
pt_update_ops->start,
pt_update_ops->last,
- XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT);
+ type);
if (IS_ERR(mjob)) {
err = PTR_ERR(mjob);
goto free_ijob;
diff --git a/drivers/gpu/drm/xe/xe_sync.c b/drivers/gpu/drm/xe/xe_sync.c
index 37866768d64c..06b1c913588a 100644
--- a/drivers/gpu/drm/xe/xe_sync.c
+++ b/drivers/gpu/drm/xe/xe_sync.c
@@ -345,15 +345,9 @@ xe_sync_in_fence_get(struct xe_sync_entry *sync, int num_sync,
return ERR_PTR(-EOPNOTSUPP);
if (q->flags & EXEC_QUEUE_FLAG_VM) {
- struct xe_exec_queue *__q;
- struct xe_tile *tile;
- u8 id;
-
- for_each_tile(tile, vm->xe, id) {
+ num_fence++;
+ for_each_tlb_inval(q, i)
num_fence++;
- for_each_tlb_inval(i)
- num_fence++;
- }
fences = kmalloc_objs(*fences, num_fence);
if (!fences)
@@ -361,17 +355,9 @@ xe_sync_in_fence_get(struct xe_sync_entry *sync, int num_sync,
fences[current_fence++] =
xe_exec_queue_last_fence_get(q, vm);
- for_each_tlb_inval(i)
+ for_each_tlb_inval(q, i)
fences[current_fence++] =
xe_exec_queue_tlb_inval_last_fence_get(q, vm, i);
- list_for_each_entry(__q, &q->multi_gt_list,
- multi_gt_link) {
- fences[current_fence++] =
- xe_exec_queue_last_fence_get(__q, vm);
- for_each_tlb_inval(i)
- fences[current_fence++] =
- xe_exec_queue_tlb_inval_last_fence_get(__q, vm, i);
- }
xe_assert(vm->xe, current_fence == num_fence);
cf = dma_fence_array_create(num_fence, fences,
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_job.c b/drivers/gpu/drm/xe/xe_tlb_inval_job.c
index 04d21015cd5d..81f560068d3c 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_job.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_job.c
@@ -39,8 +39,8 @@ struct xe_tlb_inval_job {
u64 start;
/** @end: End address to invalidate */
u64 end;
- /** @type: GT type */
- int type;
+ /** @idx: Index of tlb invalidation */
+ int idx;
/** @fence_armed: Fence has been armed */
bool fence_armed;
};
@@ -87,7 +87,7 @@ static const struct xe_dep_job_ops dep_job_ops = {
* @vm: VM which TLB invalidation is being issued for
* @start: Start address to invalidate
* @end: End address to invalidate
- * @type: GT type
+ * @idx: Index of tlb invalidation
*
* Create a TLB invalidation job and initialize internal fields. The caller is
* responsible for releasing the creation reference.
@@ -97,7 +97,7 @@ static const struct xe_dep_job_ops dep_job_ops = {
struct xe_tlb_inval_job *
xe_tlb_inval_job_create(struct xe_exec_queue *q, struct xe_tlb_inval *tlb_inval,
struct xe_dep_scheduler *dep_scheduler,
- struct xe_vm *vm, u64 start, u64 end, int type)
+ struct xe_vm *vm, u64 start, u64 end, int idx)
{
struct xe_tlb_inval_job *job;
struct drm_sched_entity *entity =
@@ -105,8 +105,7 @@ xe_tlb_inval_job_create(struct xe_exec_queue *q, struct xe_tlb_inval *tlb_inval,
struct xe_tlb_inval_fence *ifence;
int err;
- xe_assert(vm->xe, type == XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT ||
- type == XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT);
+ xe_assert(vm->xe, idx < XE_EXEC_QUEUE_TLB_INVAL_COUNT);
job = kmalloc_obj(*job);
if (!job)
@@ -120,7 +119,7 @@ xe_tlb_inval_job_create(struct xe_exec_queue *q, struct xe_tlb_inval *tlb_inval,
job->fence_armed = false;
xe_page_reclaim_list_init(&job->prl);
job->dep.ops = &dep_job_ops;
- job->type = type;
+ job->idx = idx;
kref_init(&job->refcount);
xe_exec_queue_get(q); /* Pairs with put in xe_tlb_inval_job_destroy */
xe_vm_get(vm); /* Pairs with put in xe_tlb_inval_job_destroy */
@@ -280,7 +279,7 @@ struct dma_fence *xe_tlb_inval_job_push(struct xe_tlb_inval_job *job,
/* Let the upper layers fish this out */
xe_exec_queue_tlb_inval_last_fence_set(job->q, job->vm,
&job->dep.drm.s_fence->finished,
- job->type);
+ job->idx);
xe_migrate_job_unlock(m, job->q);
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_job.h b/drivers/gpu/drm/xe/xe_tlb_inval_job.h
index 03d6e21cd611..2a4478f529e6 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_job.h
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_job.h
@@ -20,7 +20,7 @@ struct xe_vm;
struct xe_tlb_inval_job *
xe_tlb_inval_job_create(struct xe_exec_queue *q, struct xe_tlb_inval *tlb_inval,
struct xe_dep_scheduler *dep_scheduler,
- struct xe_vm *vm, u64 start, u64 end, int type);
+ struct xe_vm *vm, u64 start, u64 end, int idx);
void xe_tlb_inval_job_add_page_reclaim(struct xe_tlb_inval_job *job,
struct xe_page_reclaim_list *prl);
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index c0d8c0e266a3..feb67a565fb0 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -1816,7 +1816,7 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
struct xe_exec_queue *q;
u32 create_flags = EXEC_QUEUE_FLAG_VM;
- if (!vm->pt_root[id])
+ if (!vm->pt_root[id] || vm->q)
continue;
if (!xef) /* Not from userspace */
@@ -1827,7 +1827,7 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
err = PTR_ERR(q);
goto err_close;
}
- vm->q[id] = q;
+ vm->q = q;
}
}
@@ -1938,24 +1938,18 @@ void xe_vm_close_and_put(struct xe_vm *vm)
if (xe_vm_in_fault_mode(vm))
xe_svm_close(vm);
- down_write(&vm->lock);
- for_each_tile(tile, xe, id) {
- if (vm->q[id]) {
- int i;
+ if (vm->q) {
+ int i;
- xe_exec_queue_last_fence_put(vm->q[id], vm);
- for_each_tlb_inval(i)
- xe_exec_queue_tlb_inval_last_fence_put(vm->q[id], vm, i);
- }
- }
- up_write(&vm->lock);
+ down_write(&vm->lock);
+ xe_exec_queue_last_fence_put(vm->q, vm);
+ for_each_tlb_inval(vm->q, i)
+ xe_exec_queue_tlb_inval_last_fence_put(vm->q, vm, i);
+ up_write(&vm->lock);
- for_each_tile(tile, xe, id) {
- if (vm->q[id]) {
- xe_exec_queue_kill(vm->q[id]);
- xe_exec_queue_put(vm->q[id]);
- vm->q[id] = NULL;
- }
+ xe_exec_queue_kill(vm->q);
+ xe_exec_queue_put(vm->q);
+ vm->q = NULL;
}
down_write(&vm->lock);
@@ -2086,7 +2080,7 @@ u64 xe_vm_pdp4_descriptor(struct xe_vm *vm, struct xe_tile *tile)
static struct xe_exec_queue *
to_wait_exec_queue(struct xe_vm *vm, struct xe_exec_queue *q)
{
- return q ? q : vm->q[0];
+ return q ? q : vm->q;
}
static struct xe_user_fence *
@@ -3529,13 +3523,10 @@ static int vm_ops_setup_tile_args(struct xe_vm *vm, struct xe_vma_ops *vops)
if (vops->pt_update_ops[id].q)
continue;
- if (q) {
+ if (q)
vops->pt_update_ops[id].q = q;
- if (vm->pt_root[id] && !list_empty(&q->multi_gt_list))
- q = list_next_entry(q, multi_gt_list);
- } else {
- vops->pt_update_ops[id].q = vm->q[id];
- }
+ else
+ vops->pt_update_ops[id].q = vm->q;
}
return number_tiles;
@@ -3555,15 +3546,15 @@ static struct dma_fence *ops_execute(struct xe_vm *vm,
if (number_tiles == 0)
return ERR_PTR(-ENODATA);
- for_each_tile(tile, vm->xe, id) {
+ for_each_tile(tile, vm->xe, id)
++n_fence;
- if (!(vops->flags & XE_VMA_OPS_FLAG_SKIP_TLB_WAIT))
- for_each_tlb_inval(i)
- ++n_fence;
+ if (!(vops->flags & XE_VMA_OPS_FLAG_SKIP_TLB_WAIT)) {
+ for_each_tlb_inval(vops->pt_update_ops[0].q, i)
+ ++n_fence;
}
- fences = kmalloc_objs(*fences, n_fence);
+ fences = kcalloc(n_fence, sizeof(*fences), GFP_KERNEL);
if (!fences) {
fence = ERR_PTR(-ENOMEM);
goto err_trace;
@@ -3605,9 +3596,15 @@ static struct dma_fence *ops_execute(struct xe_vm *vm,
continue;
xe_migrate_job_lock(tile->migrate, q);
- for_each_tlb_inval(i)
- fences[current_fence++] =
- xe_exec_queue_tlb_inval_last_fence_get(q, vm, i);
+ for_each_tlb_inval(q, i) {
+ if (i >= (tile->id + 1) * XE_MAX_GT_PER_TILE ||
+ i < tile->id * XE_MAX_GT_PER_TILE)
+ continue;
+
+ fences[current_fence++] = fence ?
+ xe_exec_queue_tlb_inval_last_fence_get(q, vm, i) :
+ dma_fence_get_stub();
+ }
xe_migrate_job_unlock(tile->migrate, q);
}
@@ -4137,7 +4134,7 @@ int xe_vm_bind_ioctl(struct drm_device *dev, void *data, struct drm_file *file)
syncs_user = u64_to_user_ptr(args->syncs);
for (num_syncs = 0; num_syncs < args->num_syncs; num_syncs++) {
- struct xe_exec_queue *__q = q ?: vm->q[0];
+ struct xe_exec_queue *__q = q ?: vm->q;
err = xe_sync_entry_parse(xe, xef, &syncs[num_syncs],
&syncs_user[num_syncs],
diff --git a/drivers/gpu/drm/xe/xe_vm_types.h b/drivers/gpu/drm/xe/xe_vm_types.h
index 648031e64145..c91cb13fc4f1 100644
--- a/drivers/gpu/drm/xe/xe_vm_types.h
+++ b/drivers/gpu/drm/xe/xe_vm_types.h
@@ -262,7 +262,7 @@ struct xe_vm {
struct xe_device *xe;
/* exec queue used for (un)binding vma's */
- struct xe_exec_queue *q[XE_MAX_TILES_PER_DEVICE];
+ struct xe_exec_queue *q;
/** @lru_bulk_move: Bulk LRU move list for this VM's BOs */
struct ttm_lru_bulk_move lru_bulk_move;
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 16/24] drm/xe: Add CPU bind layer
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (14 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 15/24] drm/xe: Make bind queues operate cross-tile Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
` (12 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Himal Prasad Ghimiray, Shuicheng Lin
With CPU binds, it no longer makes sense to implement CPU bind handling
in the migrate layer, as these operations are entirely decoupled from
hardware. Introduce a dedicated CPU bind layer stored at the device
level.
Since CPU binds are tile-independent, update the PT layer to generate a
single bind job even when pages are mirrored across tiles.
This patch is large because the refactor touches multiple file / layers
and ensures functional equivalence before and after the change.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
---
v7:
- s/vm_ops_setup_tile_args/vm_ops_setup/ (Shuicheng)
- Lot's kernel doc updates (Shuicheng)
---
drivers/gpu/drm/xe/Makefile | 1 +
drivers/gpu/drm/xe/xe_cpu_bind.c | 296 +++++++++++++
drivers/gpu/drm/xe/xe_cpu_bind.h | 111 +++++
drivers/gpu/drm/xe/xe_device.c | 5 +
drivers/gpu/drm/xe/xe_device_types.h | 4 +
drivers/gpu/drm/xe/xe_exec_queue.c | 3 +-
drivers/gpu/drm/xe/xe_guc_submit.c | 46 +-
drivers/gpu/drm/xe/xe_migrate.c | 243 -----------
drivers/gpu/drm/xe/xe_migrate.h | 89 ----
drivers/gpu/drm/xe/xe_pt.c | 546 ++++++++++++------------
drivers/gpu/drm/xe/xe_pt.h | 8 +-
drivers/gpu/drm/xe/xe_pt_types.h | 14 -
drivers/gpu/drm/xe/xe_sched_job.c | 10 +-
drivers/gpu/drm/xe/xe_sched_job_types.h | 11 +-
drivers/gpu/drm/xe/xe_tlb_inval_job.c | 13 +-
drivers/gpu/drm/xe/xe_tlb_inval_job.h | 2 -
drivers/gpu/drm/xe/xe_vm.c | 157 ++-----
drivers/gpu/drm/xe/xe_vm_types.h | 10 +-
18 files changed, 808 insertions(+), 761 deletions(-)
create mode 100644 drivers/gpu/drm/xe/xe_cpu_bind.c
create mode 100644 drivers/gpu/drm/xe/xe_cpu_bind.h
diff --git a/drivers/gpu/drm/xe/Makefile b/drivers/gpu/drm/xe/Makefile
index ba4404896f2f..563c2f859361 100644
--- a/drivers/gpu/drm/xe/Makefile
+++ b/drivers/gpu/drm/xe/Makefile
@@ -35,6 +35,7 @@ $(obj)/generated/%_device_wa_oob.c $(obj)/generated/%_device_wa_oob.h: $(obj)/xe
xe-y += xe_bb.o \
xe_bo.o \
xe_bo_evict.o \
+ xe_cpu_bind.o \
xe_dep_scheduler.o \
xe_devcoredump.o \
xe_device.o \
diff --git a/drivers/gpu/drm/xe/xe_cpu_bind.c b/drivers/gpu/drm/xe/xe_cpu_bind.c
new file mode 100644
index 000000000000..318978298420
--- /dev/null
+++ b/drivers/gpu/drm/xe/xe_cpu_bind.c
@@ -0,0 +1,296 @@
+// SPDX-License-Identifier: MIT
+/*
+ * Copyright © 2026 Intel Corporation
+ */
+
+#include <drm/drm_managed.h>
+#include <linux/mutex.h>
+
+#include "xe_cpu_bind.h"
+#include "xe_device_types.h"
+#include "xe_exec_queue.h"
+#include "xe_pt.h"
+#include "xe_sched_job.h"
+#include "xe_trace_bo.h"
+#include "xe_vm.h"
+
+/**
+ * struct xe_cpu_bind - cpu_bind context.
+ */
+struct xe_cpu_bind {
+ /** @xe: Xe device */
+ struct xe_device *xe;
+ /** @q: Default exec queue used for kernel binds */
+ struct xe_exec_queue *q;
+ /** @job_mutex: Timeline mutex for @q. */
+ struct mutex job_mutex;
+};
+
+static bool is_cpu_bind_queue(struct xe_cpu_bind *cpu_bind,
+ struct xe_exec_queue *q)
+{
+ return cpu_bind->q == q;
+}
+
+static void xe_cpu_bind_fini(void *arg)
+{
+ struct xe_cpu_bind *cpu_bind = arg;
+
+ mutex_destroy(&cpu_bind->job_mutex);
+ xe_exec_queue_put(cpu_bind->q);
+}
+
+/**
+ * xe_cpu_bind_init() - Initialize a cpu_bind context
+ * @xe: &struct xe_device
+ *
+ * Return: 0 if successful, negative error code on failure
+ */
+int xe_cpu_bind_init(struct xe_device *xe)
+{
+ struct xe_cpu_bind *cpu_bind =
+ drmm_kzalloc(&xe->drm, sizeof(*cpu_bind), GFP_KERNEL);
+ struct xe_exec_queue *q;
+
+ if (!cpu_bind)
+ return -ENOMEM;
+
+ q = xe_exec_queue_create_bind(xe, xe_device_get_root_tile(xe), NULL,
+ EXEC_QUEUE_FLAG_KERNEL |
+ EXEC_QUEUE_FLAG_MIGRATE, 0);
+ if (IS_ERR(q))
+ return PTR_ERR(q);
+
+ cpu_bind->xe = xe;
+ cpu_bind->q = q;
+ xe->cpu_bind = cpu_bind;
+
+ mutex_init(&cpu_bind->job_mutex);
+
+ fs_reclaim_acquire(GFP_KERNEL);
+ might_lock(&cpu_bind->job_mutex);
+ fs_reclaim_release(GFP_KERNEL);
+
+ return devm_add_action_or_reset(cpu_bind->xe->drm.dev, xe_cpu_bind_fini,
+ cpu_bind);
+}
+
+/**
+ * xe_cpu_bind_queue() - Get the bind queue from cpu_bind context.
+ * @cpu_bind: The cpu bind context.
+ *
+ * Return: Pointer to bind queue.
+ */
+struct xe_exec_queue *xe_cpu_bind_queue(struct xe_cpu_bind *cpu_bind)
+{
+ return cpu_bind->q;
+}
+
+/**
+ * xe_cpu_bind_update_pgtables_execute() - Update a VM's PTEs via the CPU
+ * @vm: The VM being updated
+ * @tile: The tile being updated
+ * @ops: The CPU bind PT update ops
+ * @pt_op: The VM PT update op
+ * @num_ops: The number of The VM PT update ops
+ * @force_clear: Force clear operation
+ *
+ * Execute the VM PT update ops array which results in a VM's PTEs being updated
+ * via the CPU.
+ */
+void
+xe_cpu_bind_update_pgtables_execute(struct xe_vm *vm, struct xe_tile *tile,
+ const struct xe_cpu_bind_pt_update_ops *ops,
+ struct xe_vm_pgtable_update_op *pt_op,
+ u32 num_ops, bool force_clear)
+{
+ u32 j, i;
+
+ for (j = 0; j < num_ops; ++j, ++pt_op) {
+ for (i = 0; i < pt_op->num_entries; i++) {
+ const struct xe_vm_pgtable_update *update =
+ &pt_op->entries[i];
+
+ xe_assert(vm->xe, !iosys_map_is_null(&update->pt_bo->vmap));
+
+ if (pt_op->bind && !force_clear)
+ ops->populate(tile, &update->pt_bo->vmap,
+ update);
+ else
+ ops->clear(vm, tile, &update->pt_bo->vmap,
+ update);
+ }
+ }
+
+ trace_xe_vm_cpu_bind(vm);
+ xe_device_wmb(vm->xe);
+}
+
+static struct dma_fence *
+xe_cpu_bind_update_pgtables_no_job(struct xe_cpu_bind *cpu_bind,
+ struct xe_cpu_bind_pt_update *pt_update)
+{
+ const struct xe_cpu_bind_pt_update_ops *ops = pt_update->ops;
+ struct xe_vm *vm = pt_update->vops->vm;
+ struct xe_tile *tile;
+ int err, id;
+
+ if (ops->pre_commit) {
+ pt_update->job = NULL;
+ err = ops->pre_commit(pt_update);
+ if (err)
+ return ERR_PTR(err);
+ }
+
+ for_each_tile(tile, vm->xe, id) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &pt_update->vops->pt_update_ops[tile->id];
+
+ if (!pt_update_ops->pt_job_ops)
+ continue;
+
+ xe_cpu_bind_update_pgtables_execute(vm, tile, ops,
+ pt_update_ops->pt_job_ops->ops,
+ pt_update_ops->pt_job_ops->current_op,
+ false);
+ }
+
+ return dma_fence_get_stub();
+}
+
+static struct dma_fence *
+xe_cpu_bind_update_pgtables_job(struct xe_cpu_bind *cpu_bind,
+ struct xe_cpu_bind_pt_update *pt_update)
+{
+ const struct xe_cpu_bind_pt_update_ops *ops = pt_update->ops;
+ struct xe_exec_queue *q = pt_update->vops->q;
+ struct xe_device *xe = cpu_bind->xe;
+ struct xe_sched_job *job;
+ struct dma_fence *fence;
+ struct xe_tile *tile;
+ int err, id;
+ bool is_cpu_bind = is_cpu_bind_queue(cpu_bind, q);
+
+ job = xe_sched_job_create(q, NULL);
+ if (IS_ERR(job))
+ return ERR_CAST(job);
+
+ xe_assert(xe, job->is_pt_job);
+
+ if (ops->pre_commit) {
+ pt_update->job = job;
+ err = ops->pre_commit(pt_update);
+ if (err)
+ goto err_job;
+ }
+
+ if (is_cpu_bind)
+ mutex_lock(&cpu_bind->job_mutex);
+
+ job->pt_update[0].vm = pt_update->vops->vm;
+ job->pt_update[0].ops = ops;
+ for_each_tile(tile, xe, id) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &pt_update->vops->pt_update_ops[tile->id];
+
+ job->pt_update[0].pt_job_ops[tile->id] =
+ xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
+ }
+
+ xe_sched_job_arm(job);
+ fence = dma_fence_get(&job->drm.s_fence->finished);
+ xe_sched_job_push(job);
+
+ if (is_cpu_bind)
+ mutex_unlock(&cpu_bind->job_mutex);
+
+ return fence;
+
+err_job:
+ xe_sched_job_put(job);
+ return ERR_PTR(err);
+}
+
+/**
+ * xe_cpu_bind_update_pgtables() - Pipelined page-table update
+ * @cpu_bind: The cpu bind context.
+ * @pt_update: PT update arguments
+ *
+ * Perform a pipelined page-table update. The update descriptors are typically
+ * built under the same lock critical section as a call to this function. If
+ * using the default engine for the updates, they will be performed in the
+ * order they grab the job_mutex. If different engines are used, external
+ * synchronization is needed for overlapping updates to maintain page-table
+ * consistency. Note that the meaning of "overlapping" is that the updates
+ * touch the same page-table, which might be a higher-level page-directory.
+ * If no pipelining is needed, then updates may be performed by the cpu.
+ *
+ * Return: A dma_fence that, when signaled, indicates the update completion.
+ */
+struct dma_fence *
+xe_cpu_bind_update_pgtables(struct xe_cpu_bind *cpu_bind,
+ struct xe_cpu_bind_pt_update *pt_update)
+{
+ struct dma_fence *fence;
+
+ fence = xe_cpu_bind_update_pgtables_no_job(cpu_bind, pt_update);
+
+ /* -ETIME indicates a job is needed, anything else is legit error */
+ if (!IS_ERR(fence) || PTR_ERR(fence) != -ETIME)
+ return fence;
+
+ return xe_cpu_bind_update_pgtables_job(cpu_bind, pt_update);
+}
+
+/**
+ * xe_cpu_bind_job_lock() - Lock cpu_bind job lock
+ * @cpu_bind: The cpu bind context.
+ * @q: Queue associated with the operation which requires a lock
+ *
+ * Lock the cpu_bind job lock if the queue is a cpu bind queue, otherwise
+ * assert the VM's dma-resv is held (user queue's have own locking).
+ */
+void xe_cpu_bind_job_lock(struct xe_cpu_bind *cpu_bind,
+ struct xe_exec_queue *q)
+{
+ bool is_cpu_bind = is_cpu_bind_queue(cpu_bind, q);
+
+ if (is_cpu_bind)
+ mutex_lock(&cpu_bind->job_mutex);
+ else
+ xe_vm_assert_held(q->user_vm); /* User queues VM's should be locked */
+}
+
+/**
+ * xe_cpu_bind_job_unlock() - Unlock cpu_bind job lock
+ * @cpu_bind: The cpu bind context.
+ * @q: Queue associated with the operation which requires a lock
+ *
+ * Unlock the cpu_bind job lock if the queue is a cpu bind queue, otherwise
+ * assert the VM's dma-resv is held (user queue's have own locking).
+ */
+void xe_cpu_bind_job_unlock(struct xe_cpu_bind *cpu_bind,
+ struct xe_exec_queue *q)
+{
+ bool is_cpu_bind = is_cpu_bind_queue(cpu_bind, q);
+
+ if (is_cpu_bind)
+ mutex_unlock(&cpu_bind->job_mutex);
+ else
+ xe_vm_assert_held(q->user_vm); /* User queues VM's should be locked */
+}
+
+#if IS_ENABLED(CONFIG_PROVE_LOCKING)
+/**
+ * xe_cpu_bind_job_lock_assert() - Assert cpu_bind job lock held of queue
+ * @q: cpu bind queue
+ */
+void xe_cpu_bind_job_lock_assert(struct xe_exec_queue *q)
+{
+ struct xe_device *xe = gt_to_xe(q->gt);
+ struct xe_cpu_bind *cpu_bind = xe->cpu_bind;
+
+ xe_assert(xe, q == cpu_bind->q);
+ lockdep_assert_held(&cpu_bind->job_mutex);
+}
+#endif
diff --git a/drivers/gpu/drm/xe/xe_cpu_bind.h b/drivers/gpu/drm/xe/xe_cpu_bind.h
new file mode 100644
index 000000000000..69aad5edf204
--- /dev/null
+++ b/drivers/gpu/drm/xe/xe_cpu_bind.h
@@ -0,0 +1,111 @@
+/* SPDX-License-Identifier: MIT */
+/*
+ * Copyright © 2026 Intel Corporation
+ */
+
+#ifndef _XE_CPU_BIND_H_
+#define _XE_CPU_BIND_H_
+
+#include <linux/types.h>
+
+struct dma_fence;
+struct iosys_map;
+struct xe_cpu_bind;
+struct xe_cpu_bind_pt_update;
+struct xe_device;
+struct xe_tlb_inval_job;
+struct xe_tile;
+struct xe_vm;
+struct xe_vm_pgtable_update;
+struct xe_vm_pgtable_update_op;
+struct xe_vma_ops;
+
+/**
+ * struct xe_cpu_bind_pt_update_ops - Callbacks for the
+ * xe_cpu_bind_update_pgtables() function.
+ */
+struct xe_cpu_bind_pt_update_ops {
+ /**
+ * @populate: Populate a page-table with ptes.
+ * @tile: The tile for the current operation.
+ * @map: struct iosys_map into the memory to be populated.
+ * @update: Information about the PTEs to be inserted.
+ *
+ * This interface is intended to be used as a callback into the
+ * page-table system to populate shared page-tables with PTEs.
+ */
+ void (*populate)(struct xe_tile *tile, struct iosys_map *map,
+ const struct xe_vm_pgtable_update *update);
+ /**
+ * @clear: Clear a page-table's ptes.
+ * @vm: VM being updated
+ * @tile: The tile for the current operation.
+ * @map: struct iosys_map into the page-table to be cleared.
+ * @update: Information about the PTEs to be cleared.
+ *
+ * This interface is intended to be used as a callback into the
+ * page-table system to clear PTEs from shared page-tables.
+ */
+ void (*clear)(struct xe_vm *vm, struct xe_tile *tile,
+ struct iosys_map *map,
+ const struct xe_vm_pgtable_update *update);
+
+ /**
+ * @pre_commit: Callback to be called just before arming the
+ * sched_job.
+ * @pt_update: Pointer to embeddable callback argument.
+ *
+ * Return: 0 on success, negative error code on error.
+ */
+ int (*pre_commit)(struct xe_cpu_bind_pt_update *pt_update);
+};
+
+/**
+ * struct xe_cpu_bind_pt_update - Argument to the struct
+ * xe_cpu_bind_pt_update_ops callbacks.
+ *
+ * Intended to be subclassed to support additional arguments if necessary.
+ */
+struct xe_cpu_bind_pt_update {
+ /** @ops: Pointer to the struct xe_cpu_bind_pt_update_ops callbacks */
+ const struct xe_cpu_bind_pt_update_ops *ops;
+ /** @vops: VMA operations */
+ struct xe_vma_ops *vops;
+ /** @job: The bind job, or NULL if the update is issued immediatel */
+ struct xe_sched_job *job;
+ /**
+ * @ijobs: The TLB invalidation jobs, individual instances can be NULL
+ */
+#define XE_CPU_BIND_INVAL_JOB_COUNT 4
+ struct xe_tlb_inval_job *ijobs[XE_CPU_BIND_INVAL_JOB_COUNT];
+};
+
+int xe_cpu_bind_init(struct xe_device *xe);
+
+struct xe_exec_queue *xe_cpu_bind_queue(struct xe_cpu_bind *cpu_bind);
+
+void
+xe_cpu_bind_update_pgtables_execute(struct xe_vm *vm, struct xe_tile *tile,
+ const struct xe_cpu_bind_pt_update_ops *ops,
+ struct xe_vm_pgtable_update_op *pt_op,
+ u32 num_ops, bool force_clear);
+
+struct dma_fence *
+xe_cpu_bind_update_pgtables(struct xe_cpu_bind *cpu_bind,
+ struct xe_cpu_bind_pt_update *pt_update);
+
+void xe_cpu_bind_job_lock(struct xe_cpu_bind *cpu_bind,
+ struct xe_exec_queue *q);
+
+void xe_cpu_bind_job_unlock(struct xe_cpu_bind *cpu_bind,
+ struct xe_exec_queue *q);
+
+#if IS_ENABLED(CONFIG_PROVE_LOCKING)
+void xe_cpu_bind_job_lock_assert(struct xe_exec_queue *q);
+#else
+static inline void xe_cpu_bind_job_lock_assert(struct xe_exec_queue *q)
+{
+}
+#endif
+
+#endif
diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
index 205cb4e7f9e8..44975e7823be 100644
--- a/drivers/gpu/drm/xe/xe_device.c
+++ b/drivers/gpu/drm/xe/xe_device.c
@@ -27,6 +27,7 @@
#include "xe_bo_evict.h"
#include "xe_configfs.h"
#include "xe_debugfs.h"
+#include "xe_cpu_bind.h"
#include "xe_defaults.h"
#include "xe_devcoredump.h"
#include "xe_device_sysfs.h"
@@ -993,6 +994,10 @@ int xe_device_probe(struct xe_device *xe)
if (err)
return err;
+ err = xe_cpu_bind_init(xe);
+ if (err)
+ return err;
+
err = xe_pagefault_init(xe);
if (err)
return err;
diff --git a/drivers/gpu/drm/xe/xe_device_types.h b/drivers/gpu/drm/xe/xe_device_types.h
index 3e1d3a143def..aff2e5c8b656 100644
--- a/drivers/gpu/drm/xe/xe_device_types.h
+++ b/drivers/gpu/drm/xe/xe_device_types.h
@@ -38,6 +38,7 @@
struct drm_pagemap_shrinker;
struct intel_display;
struct intel_dg_nvm_dev;
+struct xe_cpu_bind;
struct xe_ggtt;
struct xe_i2c;
struct xe_mmio_gem;
@@ -577,6 +578,9 @@ struct xe_device {
/** @sc: System Controller */
struct xe_sysctrl sc;
+ /** @cpu_bind: CPU bind object */
+ struct xe_cpu_bind *cpu_bind;
+
/** @atomic_svm_timeslice_ms: Atomic SVM fault timeslice MS */
u32 atomic_svm_timeslice_ms;
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.c b/drivers/gpu/drm/xe/xe_exec_queue.c
index d2270f32e35d..f18229f65ad2 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.c
+++ b/drivers/gpu/drm/xe/xe_exec_queue.c
@@ -14,6 +14,7 @@
#include <uapi/drm/xe_drm.h>
#include "xe_bo.h"
+#include "xe_cpu_bind.h"
#include "xe_dep_scheduler.h"
#include "xe_device.h"
#include "xe_gt.h"
@@ -1666,7 +1667,7 @@ static void xe_exec_queue_last_fence_lockdep_assert(struct xe_exec_queue *q,
struct xe_vm *vm)
{
if (q->flags & EXEC_QUEUE_FLAG_MIGRATE) {
- xe_migrate_job_lock_assert(q);
+ xe_cpu_bind_job_lock_assert(q);
} else if (q->flags & EXEC_QUEUE_FLAG_VM) {
lockdep_assert_held(&vm->lock);
} else {
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index 9faddb7c6407..a2784763e72b 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -19,6 +19,7 @@
#include "abi/guc_klvs_abi.h"
#include "xe_assert.h"
#include "xe_bo.h"
+#include "xe_cpu_bind.h"
#include "xe_devcoredump.h"
#include "xe_device.h"
#include "xe_exec_queue.h"
@@ -39,7 +40,6 @@
#include "xe_lrc.h"
#include "xe_macros.h"
#include "xe_map.h"
-#include "xe_migrate.h"
#include "xe_mocs.h"
#include "xe_module.h"
#include "xe_pm.h"
@@ -1245,14 +1245,38 @@ static bool is_pt_job(struct xe_sched_job *job)
return job->is_pt_job;
}
-static void run_pt_job(struct xe_sched_job *job, bool force_clear)
+static void run_pt_job(struct xe_device *xe, struct xe_sched_job *job,
+ bool force_clear)
{
- xe_migrate_update_pgtables_cpu_execute(job->pt_update[0].vm,
- job->pt_update[0].tile,
- job->pt_update[0].ops,
- job->pt_update[0].pt_job_ops->ops,
- job->pt_update[0].pt_job_ops->current_op,
- force_clear);
+ struct xe_tile *tile;
+ int id;
+
+ for_each_tile(tile, xe, id) {
+ struct xe_pt_job_ops *pt_job_ops =
+ job->pt_update[0].pt_job_ops[id];
+
+ if (!pt_job_ops || !pt_job_ops->current_op)
+ continue;
+
+ xe_cpu_bind_update_pgtables_execute(job->pt_update[0].vm, tile,
+ job->pt_update[0].ops,
+ pt_job_ops->ops,
+ pt_job_ops->current_op,
+ force_clear);
+ }
+}
+
+static void put_pt_job(struct xe_device *xe, struct xe_sched_job *job)
+{
+ struct xe_tile *tile;
+ int id;
+
+ for_each_tile(tile, xe, id) {
+ struct xe_pt_job_ops *pt_job_ops =
+ job->pt_update[0].pt_job_ops[id];
+
+ xe_pt_job_ops_put(pt_job_ops);
+ }
}
static struct dma_fence *
@@ -1272,8 +1296,10 @@ guc_exec_queue_run_job(struct drm_sched_job *drm_job)
if (is_pt_job(job)) {
xe_gt_assert(guc_to_gt(guc), !exec_queue_registered(q));
- run_pt_job(job, killed_or_banned_or_wedged_or_error);
- xe_pt_job_ops_put(job->pt_update[0].pt_job_ops);
+
+ run_pt_job(guc_to_xe(guc), job,
+ killed_or_banned_or_wedged_or_error);
+ put_pt_job(guc_to_xe(guc), job);
dma_fence_put(job->fence); /* Drop ref from xe_sched_job_arm */
return NULL;
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 1c16d31f0d10..f7e1a81434b2 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -51,8 +51,6 @@
struct xe_migrate {
/** @q: Default exec queue used for migration */
struct xe_exec_queue *q;
- /** @bind_q: Default exec queue used for binds */
- struct xe_exec_queue *bind_q;
/** @tile: Backpointer to the tile this struct xe_migrate belongs to. */
struct xe_tile *tile;
/** @job_mutex: Timeline mutex for @eng. */
@@ -110,7 +108,6 @@ static void xe_migrate_fini(void *arg)
mutex_destroy(&m->job_mutex);
xe_vm_close_and_put(m->q->vm);
xe_exec_queue_put(m->q);
- xe_exec_queue_put(m->bind_q);
}
static inline u16 xe_migrate_pat_index(struct xe_device *xe,
@@ -483,15 +480,6 @@ int xe_migrate_init(struct xe_migrate *m)
goto err_out;
}
- m->bind_q = xe_exec_queue_create(xe, vm, logical_mask, 1, hwe0,
- EXEC_QUEUE_FLAG_KERNEL |
- EXEC_QUEUE_FLAG_HIGH_PRIORITY |
- EXEC_QUEUE_FLAG_MIGRATE, 0);
- if (IS_ERR(m->bind_q)) {
- err = PTR_ERR(m->bind_q);
- goto err_out;
- }
-
/*
* XXX: Currently only reserving 1 (likely slow) BCS instance on
* PVC, may want to revisit if performance is needed.
@@ -502,15 +490,6 @@ int xe_migrate_init(struct xe_migrate *m)
EXEC_QUEUE_FLAG_MIGRATE |
EXEC_QUEUE_FLAG_LOW_LATENCY, 0);
} else {
- m->bind_q = xe_exec_queue_create_class(xe, primary_gt, vm,
- XE_ENGINE_CLASS_COPY,
- EXEC_QUEUE_FLAG_KERNEL |
- EXEC_QUEUE_FLAG_MIGRATE, 0);
- if (IS_ERR(m->bind_q)) {
- err = PTR_ERR(m->bind_q);
- goto err_out;
- }
-
m->q = xe_exec_queue_create_class(xe, primary_gt, vm,
XE_ENGINE_CLASS_COPY,
EXEC_QUEUE_FLAG_KERNEL |
@@ -546,8 +525,6 @@ int xe_migrate_init(struct xe_migrate *m)
return err;
err_out:
- if (!IS_ERR_OR_NULL(m->bind_q))
- xe_exec_queue_put(m->bind_q);
xe_vm_close_and_put(vm);
return err;
@@ -1507,17 +1484,6 @@ static u32 blt_mem_set_cmd_len(struct xe_device *xe)
return 7;
}
-/**
- * xe_get_migrate_bind_queue() - Get the bind queue from migrate context.
- * @migrate: Migrate context.
- *
- * Return: Pointer to bind queue on success, error on failure
- */
-struct xe_exec_queue *xe_migrate_bind_queue(struct xe_migrate *migrate)
-{
- return migrate->bind_q;
-}
-
static void emit_clear_link_copy(struct xe_gt *gt, struct xe_bb *bb, u64 src_ofs,
u32 size, u32 pitch)
{
@@ -1782,165 +1748,6 @@ struct migrate_test_params {
container_of(_priv, struct migrate_test_params, base)
#endif
-/**
- * xe_migrate_update_pgtables_cpu_execute() - Update a VM's PTEs via the CPU
- * @vm: The VM being updated
- * @tile: The tile being updated
- * @ops: The migrate PT update ops
- * @pt_ops: The VM PT update ops
- * @num_ops: The number of The VM PT update ops
- * @force_clear: Force clear
- *
- * Execute the VM PT update ops array which results in a VM's PTEs being updated
- * via the CPU.
- */
-void
-xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
- const struct xe_migrate_pt_update_ops *ops,
- struct xe_vm_pgtable_update_op *pt_op,
- u32 num_ops, bool force_clear)
-{
- u32 j, i;
-
- for (j = 0; j < num_ops; ++j, ++pt_op) {
- for (i = 0; i < pt_op->num_entries; i++) {
- const struct xe_vm_pgtable_update *update =
- &pt_op->entries[i];
-
- xe_tile_assert(tile, !iosys_map_is_null(&update->pt_bo->vmap));
-
- if (pt_op->bind && !force_clear)
- ops->populate(tile, &update->pt_bo->vmap,
- update);
- else
- ops->clear(vm, tile, &update->pt_bo->vmap,
- update);
- }
- }
-
- trace_xe_vm_cpu_bind(vm);
- xe_device_wmb(vm->xe);
-}
-
-static struct dma_fence *
-xe_migrate_update_pgtables_cpu(struct xe_migrate *m,
- struct xe_migrate_pt_update *pt_update)
-{
- XE_TEST_DECLARE(struct migrate_test_params *test =
- to_migrate_test_params
- (xe_cur_kunit_priv(XE_TEST_LIVE_MIGRATE));)
- const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
- struct xe_vm *vm = pt_update->vops->vm;
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &pt_update->vops->pt_update_ops[pt_update->tile_id];
- int err;
-
- if (XE_TEST_ONLY(test && test->force_gpu))
- return ERR_PTR(-ETIME);
-
- if (ops->pre_commit) {
- pt_update->job = NULL;
- err = ops->pre_commit(pt_update);
- if (err)
- return ERR_PTR(err);
- }
-
- xe_migrate_update_pgtables_cpu_execute(vm, m->tile, ops,
- pt_update_ops->pt_job_ops->ops,
- pt_update_ops->num_ops, false);
-
- return dma_fence_get_stub();
-}
-
-static bool is_migrate_queue(struct xe_migrate *m, struct xe_exec_queue *q)
-{
- return m->bind_q == q;
-}
-
-static struct dma_fence *
-__xe_migrate_update_pgtables(struct xe_migrate *m,
- struct xe_migrate_pt_update *pt_update,
- struct xe_vm_pgtable_update_ops *pt_update_ops)
-{
- const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
- struct xe_tile *tile = m->tile;
- struct xe_sched_job *job;
- struct dma_fence *fence;
- bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
- int err;
-
- job = xe_sched_job_create(pt_update_ops->q, NULL);
- if (IS_ERR(job)) {
- err = PTR_ERR(job);
- goto err_out;
- }
-
- xe_tile_assert(tile, job->is_pt_job);
-
- if (ops->pre_commit) {
- pt_update->job = job;
- err = ops->pre_commit(pt_update);
- if (err)
- goto err_job;
- }
- if (is_migrate)
- mutex_lock(&m->job_mutex);
-
- job->pt_update[0].vm = pt_update->vops->vm;
- job->pt_update[0].tile = tile;
- job->pt_update[0].ops = ops;
- job->pt_update[0].pt_job_ops =
- xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
-
- xe_sched_job_arm(job);
- fence = dma_fence_get(&job->drm.s_fence->finished);
- xe_sched_job_push(job);
-
- if (is_migrate)
- mutex_unlock(&m->job_mutex);
-
- return fence;
-
-err_job:
- xe_sched_job_put(job);
-err_out:
- return ERR_PTR(err);
-}
-
-/**
- * xe_migrate_update_pgtables() - Pipelined page-table update
- * @m: The migrate context.
- * @pt_update: PT update arguments
- *
- * Perform a pipelined page-table update. The update descriptors are typically
- * built under the same lock critical section as a call to this function. If
- * using the default engine for the updates, they will be performed in the
- * order they grab the job_mutex. If different engines are used, external
- * synchronization is needed for overlapping updates to maintain page-table
- * consistency. Note that the meaning of "overlapping" is that the updates
- * touch the same page-table, which might be a higher-level page-directory.
- * If no pipelining is needed, then updates may be performed by the cpu.
- *
- * Return: A dma_fence that, when signaled, indicates the update completion.
- */
-struct dma_fence *
-xe_migrate_update_pgtables(struct xe_migrate *m,
- struct xe_migrate_pt_update *pt_update)
-
-{
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &pt_update->vops->pt_update_ops[pt_update->tile_id];
- struct dma_fence *fence;
-
- fence = xe_migrate_update_pgtables_cpu(m, pt_update);
-
- /* -ETIME indicates a job is needed, anything else is legit error */
- if (!IS_ERR(fence) || PTR_ERR(fence) != -ETIME)
- return fence;
-
- return __xe_migrate_update_pgtables(m, pt_update, pt_update_ops);
-}
-
/**
* xe_migrate_wait() - Complete all operations using the xe_migrate context
* @m: Migrate context to wait for.
@@ -2442,56 +2249,6 @@ int xe_migrate_access_memory(struct xe_migrate *m, struct xe_bo *bo,
return IS_ERR(fence) ? PTR_ERR(fence) : 0;
}
-/**
- * xe_migrate_job_lock() - Lock migrate job lock
- * @m: The migration context.
- * @q: Queue associated with the operation which requires a lock
- *
- * Lock the migrate job lock if the queue is a migration queue, otherwise
- * assert the VM's dma-resv is held (user queue's have own locking).
- */
-void xe_migrate_job_lock(struct xe_migrate *m, struct xe_exec_queue *q)
-{
- bool is_migrate = is_migrate_queue(m, q);
-
- if (is_migrate)
- mutex_lock(&m->job_mutex);
- else
- xe_vm_assert_held(q->user_vm); /* User queues VM's should be locked */
-}
-
-/**
- * xe_migrate_job_unlock() - Unlock migrate job lock
- * @m: The migration context.
- * @q: Queue associated with the operation which requires a lock
- *
- * Unlock the migrate job lock if the queue is a migration queue, otherwise
- * assert the VM's dma-resv is held (user queue's have own locking).
- */
-void xe_migrate_job_unlock(struct xe_migrate *m, struct xe_exec_queue *q)
-{
- bool is_migrate = is_migrate_queue(m, q);
-
- if (is_migrate)
- mutex_unlock(&m->job_mutex);
- else
- xe_vm_assert_held(q->user_vm); /* User queues VM's should be locked */
-}
-
-#if IS_ENABLED(CONFIG_PROVE_LOCKING)
-/**
- * xe_migrate_job_lock_assert() - Assert migrate job lock held of queue
- * @q: Migrate queue
- */
-void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
-{
- struct xe_migrate *m = gt_to_tile(q->gt)->migrate;
-
- xe_gt_assert(q->gt, q == m->bind_q);
- lockdep_assert_held(&m->job_mutex);
-}
-#endif
-
#if IS_ENABLED(CONFIG_DRM_XE_KUNIT_TEST)
#include "tests/xe_migrate.c"
#endif
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index 94b8d8d63720..fa381ec36ef1 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -34,73 +34,6 @@ enum xe_migrate_copy_dir {
XE_MIGRATE_COPY_TO_SRAM,
};
-/**
- * struct xe_migrate_pt_update_ops - Callbacks for the
- * xe_migrate_update_pgtables() function.
- */
-struct xe_migrate_pt_update_ops {
- /**
- * @populate: Populate a command buffer or page-table with ptes.
- * @tile: The tile for the current operation.
- * @map: struct iosys_map into the memory to be populated.
- * @update: Information about the PTEs to be inserted.
- *
- * This interface is intended to be used as a callback into the
- * page-table system to populate command buffers or shared
- * page-tables with PTEs.
- */
- void (*populate)(struct xe_tile *tile, struct iosys_map *map,
- const struct xe_vm_pgtable_update *update);
- /**
- * @clear: Clear a command buffer or page-table with ptes.
- * @vm: VM being updated
- * @tile: The tile for the current operation.
- * @map: struct iosys_map into the memory to be populated.
- * @update: Information about the PTEs to be inserted.
- *
- * This interface is intended to be used as a callback into the
- * page-table system to populate command buffers or shared
- * page-tables with PTEs.
- */
- void (*clear)(struct xe_vm *vm, struct xe_tile *tile,
- struct iosys_map *map,
- const struct xe_vm_pgtable_update *update);
-
- /**
- * @pre_commit: Callback to be called just before arming the
- * sched_job.
- * @pt_update: Pointer to embeddable callback argument.
- *
- * Return: 0 on success, negative error code on error.
- */
- int (*pre_commit)(struct xe_migrate_pt_update *pt_update);
-};
-
-/**
- * struct xe_migrate_pt_update - Argument to the
- * struct xe_migrate_pt_update_ops callbacks.
- *
- * Intended to be subclassed to support additional arguments if necessary.
- */
-struct xe_migrate_pt_update {
- /** @ops: Pointer to the struct xe_migrate_pt_update_ops callbacks */
- const struct xe_migrate_pt_update_ops *ops;
- /** @vops: VMA operations */
- struct xe_vma_ops *vops;
- /** @job: The job if a GPU page-table update. NULL otherwise */
- struct xe_sched_job *job;
- /**
- * @ijob: The TLB invalidation job for primary GT. NULL otherwise
- */
- struct xe_tlb_inval_job *ijob;
- /**
- * @mjob: The TLB invalidation job for media GT. NULL otherwise
- */
- struct xe_tlb_inval_job *mjob;
- /** @tile_id: Tile ID of the update */
- u8 tile_id;
-};
-
struct xe_migrate *xe_migrate_alloc(struct xe_tile *tile);
int xe_migrate_init(struct xe_migrate *m);
@@ -138,7 +71,6 @@ void xe_migrate_ccs_rw_copy_clear(struct xe_bo *src_bo,
struct xe_lrc *xe_migrate_lrc(struct xe_migrate *migrate);
struct xe_exec_queue *xe_migrate_exec_queue(struct xe_migrate *migrate);
-struct xe_exec_queue *xe_migrate_bind_queue(struct xe_migrate *migrate);
struct dma_fence *xe_migrate_vram_copy_chunk(struct xe_bo *vram_bo, u64 vram_offset,
struct xe_bo *sysmem_bo, u64 sysmem_offset,
u64 size, enum xe_migrate_copy_dir dir);
@@ -157,29 +89,8 @@ struct dma_fence *xe_migrate_clear(struct xe_migrate *m,
struct xe_vm *xe_migrate_get_vm(struct xe_migrate *m);
-void
-xe_migrate_update_pgtables_cpu_execute(struct xe_vm *vm, struct xe_tile *tile,
- const struct xe_migrate_pt_update_ops *ops,
- struct xe_vm_pgtable_update_op *pt_op,
- u32 num_ops, bool force_clear);
-
-struct dma_fence *
-xe_migrate_update_pgtables(struct xe_migrate *m,
- struct xe_migrate_pt_update *pt_update);
-
void xe_migrate_wait(struct xe_migrate *m);
-#if IS_ENABLED(CONFIG_PROVE_LOCKING)
-void xe_migrate_job_lock_assert(struct xe_exec_queue *q);
-#else
-static inline void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
-{
-}
-#endif
-
-void xe_migrate_job_lock(struct xe_migrate *m, struct xe_exec_queue *q);
-void xe_migrate_job_unlock(struct xe_migrate *m, struct xe_exec_queue *q);
-
#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
struct xe_bo *scratch_bo,
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index bf0168d071c0..fbefcfc52dca 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -7,12 +7,12 @@
#include "regs/xe_gtt_defs.h"
#include "xe_bo.h"
+#include "xe_cpu_bind.h"
#include "xe_device.h"
#include "xe_drm_client.h"
#include "xe_exec_queue.h"
#include "xe_gt.h"
#include "xe_gt_stats.h"
-#include "xe_migrate.h"
#include "xe_page_reclaim.h"
#include "xe_pat.h"
#include "xe_pt_types.h"
@@ -1400,11 +1400,9 @@ static int op_add_deps(struct xe_vm *vm, struct xe_vma_op *op,
}
static int xe_pt_vm_dependencies(struct xe_sched_job *job,
- struct xe_tlb_inval_job *ijob,
- struct xe_tlb_inval_job *mjob,
+ struct xe_tlb_inval_job **ijobs,
struct xe_vm *vm,
struct xe_vma_ops *vops,
- struct xe_vm_pgtable_update_ops *pt_update_ops,
struct xe_range_fence_tree *rftree)
{
struct xe_range_fence *rtfence;
@@ -1417,20 +1415,22 @@ static int xe_pt_vm_dependencies(struct xe_sched_job *job,
if (!job && !no_in_syncs(vops->syncs, vops->num_syncs))
return -ETIME;
- if (!job && !xe_exec_queue_is_idle(pt_update_ops->q))
+ if (!job && !xe_exec_queue_is_idle(vops->q))
return -ETIME;
- if (pt_update_ops->wait_vm_bookkeep || pt_update_ops->wait_vm_kernel) {
- err = job_test_add_deps(job, xe_vm_resv(vm),
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_BOOKKEEP :
- DMA_RESV_USAGE_KERNEL);
+ if (vops->flags & (XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP |
+ XE_VMA_OPS_FLAG_WAIT_VM_KERNEL)) {
+ enum dma_resv_usage usage = DMA_RESV_USAGE_KERNEL;
+
+ if (vops->flags & XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP)
+ usage = DMA_RESV_USAGE_BOOKKEEP;
+
+ err = job_test_add_deps(job, xe_vm_resv(vm), usage);
if (err)
return err;
}
- rtfence = xe_range_fence_tree_first(rftree, pt_update_ops->start,
- pt_update_ops->last);
+ rtfence = xe_range_fence_tree_first(rftree, vops->start, vops->last);
while (rtfence) {
fence = rtfence->fence;
@@ -1448,9 +1448,8 @@ static int xe_pt_vm_dependencies(struct xe_sched_job *job,
return err;
}
- rtfence = xe_range_fence_tree_next(rtfence,
- pt_update_ops->start,
- pt_update_ops->last);
+ rtfence = xe_range_fence_tree_next(rtfence, vops->start,
+ vops->last);
}
list_for_each_entry(op, &vops->list, link) {
@@ -1463,14 +1462,11 @@ static int xe_pt_vm_dependencies(struct xe_sched_job *job,
err = xe_sync_entry_add_deps(&vops->syncs[i], job);
if (job) {
- if (ijob) {
- err = xe_tlb_inval_job_alloc_dep(ijob);
- if (err)
- return err;
- }
+ for (i = 0; i < XE_CPU_BIND_INVAL_JOB_COUNT; ++i) {
+ if (!ijobs[i])
+ continue;
- if (mjob) {
- err = xe_tlb_inval_job_alloc_dep(mjob);
+ err = xe_tlb_inval_job_alloc_dep(ijobs[i]);
if (err)
return err;
}
@@ -1479,17 +1475,14 @@ static int xe_pt_vm_dependencies(struct xe_sched_job *job,
return err;
}
-static int xe_pt_pre_commit(struct xe_migrate_pt_update *pt_update)
+static int xe_pt_pre_commit(struct xe_cpu_bind_pt_update *pt_update)
{
struct xe_vma_ops *vops = pt_update->vops;
struct xe_vm *vm = vops->vm;
- struct xe_range_fence_tree *rftree = &vm->rftree[pt_update->tile_id];
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &vops->pt_update_ops[pt_update->tile_id];
+ struct xe_range_fence_tree *rftree = &vm->rftree;
- return xe_pt_vm_dependencies(pt_update->job, pt_update->ijob,
- pt_update->mjob, vm, pt_update->vops,
- pt_update_ops, rftree);
+ return xe_pt_vm_dependencies(pt_update->job, pt_update->ijobs,
+ vm, vops, rftree);
}
#if IS_ENABLED(CONFIG_DRM_GPUSVM)
@@ -1549,8 +1542,7 @@ static bool xe_pt_userptr_inject_eagain(struct xe_userptr_vma *uvma)
#endif
-static int vma_check_userptr(struct xe_vm *vm, struct xe_vma *vma,
- struct xe_vm_pgtable_update_ops *pt_update)
+static int vma_check_userptr(struct xe_vm *vm, struct xe_vma *vma)
{
struct xe_userptr_vma *uvma;
unsigned long notifier_seq;
@@ -1580,8 +1572,7 @@ static int vma_check_userptr(struct xe_vm *vm, struct xe_vma *vma,
return 0;
}
-static int op_check_svm_userptr(struct xe_vm *vm, struct xe_vma_op *op,
- struct xe_vm_pgtable_update_ops *pt_update)
+static int op_check_svm_userptr(struct xe_vm *vm, struct xe_vma_op *op)
{
int err = 0;
@@ -1592,13 +1583,13 @@ static int op_check_svm_userptr(struct xe_vm *vm, struct xe_vma_op *op,
if (!op->map.immediate && xe_vm_in_fault_mode(vm))
break;
- err = vma_check_userptr(vm, op->map.vma, pt_update);
+ err = vma_check_userptr(vm, op->map.vma);
break;
case DRM_GPUVA_OP_REMAP:
if (op->remap.prev && !op->remap.skip_prev)
- err = vma_check_userptr(vm, op->remap.prev, pt_update);
+ err = vma_check_userptr(vm, op->remap.prev);
if (!err && op->remap.next && !op->remap.skip_next)
- err = vma_check_userptr(vm, op->remap.next, pt_update);
+ err = vma_check_userptr(vm, op->remap.next);
break;
case DRM_GPUVA_OP_UNMAP:
break;
@@ -1618,7 +1609,7 @@ static int op_check_svm_userptr(struct xe_vm *vm, struct xe_vma_op *op,
}
}
} else {
- err = vma_check_userptr(vm, gpuva_to_vma(op->base.prefetch.va), pt_update);
+ err = vma_check_userptr(vm, gpuva_to_vma(op->base.prefetch.va));
}
break;
#if IS_ENABLED(CONFIG_DRM_XE_GPUSVM)
@@ -1644,12 +1635,10 @@ static int op_check_svm_userptr(struct xe_vm *vm, struct xe_vma_op *op,
return err;
}
-static int xe_pt_svm_userptr_pre_commit(struct xe_migrate_pt_update *pt_update)
+static int xe_pt_svm_userptr_pre_commit(struct xe_cpu_bind_pt_update *pt_update)
{
struct xe_vm *vm = pt_update->vops->vm;
struct xe_vma_ops *vops = pt_update->vops;
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &vops->pt_update_ops[pt_update->tile_id];
struct xe_vma_op *op;
int err;
@@ -1660,7 +1649,7 @@ static int xe_pt_svm_userptr_pre_commit(struct xe_migrate_pt_update *pt_update)
xe_pt_svm_userptr_notifier_lock(vm);
list_for_each_entry(op, &vops->list, link) {
- err = op_check_svm_userptr(vm, op, pt_update_ops);
+ err = op_check_svm_userptr(vm, op);
if (err) {
xe_pt_svm_userptr_notifier_unlock(vm);
break;
@@ -2007,9 +1996,9 @@ static unsigned int xe_pt_stage_unbind(struct xe_tile *tile,
}
static void
-xe_migrate_clear_pgtable_callback(struct xe_vm *vm, struct xe_tile *tile,
- struct iosys_map *map,
- const struct xe_vm_pgtable_update *update)
+xe_pt_clear_pgtable_callback(struct xe_vm *vm, struct xe_tile *tile,
+ struct iosys_map *map,
+ const struct xe_vm_pgtable_update *update)
{
u64 empty = __xe_pt_empty_pte(tile, vm, update->level);
int i;
@@ -2087,6 +2076,9 @@ to_pt_op(struct xe_vm_pgtable_update_ops *pt_update_ops, u32 op_idx)
static u32
get_current_op(struct xe_vm_pgtable_update_ops *pt_update_ops)
{
+ if (!pt_update_ops->pt_job_ops)
+ return 0;
+
return pt_update_ops->pt_job_ops->current_op;
}
@@ -2376,6 +2368,7 @@ static int unbind_range_prepare(struct xe_vm *vm,
static int op_prepare(struct xe_vm *vm,
struct xe_tile *tile,
+ struct xe_vma_ops *vops,
struct xe_vm_pgtable_update_ops *pt_update_ops,
struct xe_vma_op *op)
{
@@ -2392,7 +2385,7 @@ static int op_prepare(struct xe_vm *vm,
err = bind_op_prepare(vm, tile, pt_update_ops, op->map.vma,
op->map.invalidate_on_bind);
- pt_update_ops->wait_vm_kernel = true;
+ vops->flags |= XE_VMA_OPS_FLAG_WAIT_VM_KERNEL;
break;
case DRM_GPUVA_OP_REMAP:
{
@@ -2406,12 +2399,12 @@ static int op_prepare(struct xe_vm *vm,
if (!err && op->remap.prev && !op->remap.skip_prev) {
err = bind_op_prepare(vm, tile, pt_update_ops,
op->remap.prev, false);
- pt_update_ops->wait_vm_bookkeep = true;
+ vops->flags |= XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
}
if (!err && op->remap.next && !op->remap.skip_next) {
err = bind_op_prepare(vm, tile, pt_update_ops,
op->remap.next, false);
- pt_update_ops->wait_vm_bookkeep = true;
+ vops->flags |= XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
}
break;
}
@@ -2447,7 +2440,7 @@ static int op_prepare(struct xe_vm *vm,
}
} else {
err = bind_op_prepare(vm, tile, pt_update_ops, vma, false);
- pt_update_ops->wait_vm_kernel = true;
+ vops->flags |= XE_VMA_OPS_FLAG_WAIT_VM_KERNEL;
}
break;
}
@@ -2481,18 +2474,8 @@ xe_pt_update_ops_init(struct xe_vm_pgtable_update_ops *pt_update_ops)
xe_page_reclaim_list_init(&pt_update_ops->prl);
}
-/**
- * xe_pt_update_ops_prepare() - Prepare PT update operations
- * @tile: Tile of PT update operations
- * @vops: VMA operationa
- *
- * Prepare PT update operations which includes updating internal PT state,
- * allocate memory for page tables, populate page table being pruned in, and
- * create PT update operations for leaf insertion / removal.
- *
- * Return: 0 on success, negative error code on error.
- */
-int xe_pt_update_ops_prepare(struct xe_tile *tile, struct xe_vma_ops *vops)
+static int __xe_pt_update_ops_prepare(struct xe_tile *tile,
+ struct xe_vma_ops *vops)
{
struct xe_vm_pgtable_update_ops *pt_update_ops =
&vops->pt_update_ops[tile->id];
@@ -2511,7 +2494,7 @@ int xe_pt_update_ops_prepare(struct xe_tile *tile, struct xe_vma_ops *vops)
return err;
list_for_each_entry(op, &vops->list, link) {
- err = op_prepare(vops->vm, tile, pt_update_ops, op);
+ err = op_prepare(vops->vm, tile, vops, pt_update_ops, op);
if (err)
return err;
@@ -2520,6 +2503,16 @@ int xe_pt_update_ops_prepare(struct xe_tile *tile, struct xe_vma_ops *vops)
xe_tile_assert(tile, get_current_op(pt_update_ops) <=
pt_update_ops->num_ops);
+ /* Propagate individual tile state up to VMA operation */
+ if (pt_update_ops->start < vops->start)
+ vops->start = pt_update_ops->start;
+ if (pt_update_ops->last > vops->last)
+ vops->last = pt_update_ops->last;
+ if (pt_update_ops->needs_invalidation)
+ vops->flags |= XE_VMA_OPS_FLAG_NEEDS_INVALIDATION;
+ if (pt_update_ops->needs_svm_lock)
+ vops->flags |= XE_VMA_OPS_FLAG_NEEDS_SVM_LOCK;
+
#ifdef TEST_VM_OPS_ERROR
if (vops->inject_error &&
vops->vm->xe->vm_inject_error_position == FORCE_OP_ERROR_PREPARE)
@@ -2528,35 +2521,68 @@ int xe_pt_update_ops_prepare(struct xe_tile *tile, struct xe_vma_ops *vops)
return 0;
}
-ALLOW_ERROR_INJECTION(xe_pt_update_ops_prepare, ERRNO);
-static void bind_op_commit(struct xe_vm *vm, struct xe_tile *tile,
- struct xe_vm_pgtable_update_ops *pt_update_ops,
- struct xe_vma *vma, struct dma_fence *fence,
- struct dma_fence *fence2, bool invalidate_on_bind)
+/**
+ * xe_pt_update_ops_prepare() - Prepare PT update operations
+ * @xe: xe device.
+ * @vops: VMA operationa
+ *
+ * Prepare PT update operations which includes updating internal PT state,
+ * allocate memory for page tables, populate page table being pruned in, and
+ * create PT update operations for leaf insertion / removal.
+ *
+ * Return: 0 on success, negative error code on error.
+ */
+int xe_pt_update_ops_prepare(struct xe_device *xe, struct xe_vma_ops *vops)
{
- xe_tile_assert(tile, !xe_vma_is_cpu_addr_mirror(vma));
+ struct xe_tile *tile;
+ int id, err;
+
+ for_each_tile(tile, xe, id) {
+ if (!vops->pt_update_ops[id].num_ops)
+ continue;
- if (!xe_vma_has_no_bo(vma) && !xe_vma_bo(vma)->vm) {
- dma_resv_add_fence(xe_vma_bo(vma)->ttm.base.resv, fence,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
- if (fence2)
- dma_resv_add_fence(xe_vma_bo(vma)->ttm.base.resv, fence2,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
+ err = __xe_pt_update_ops_prepare(tile, vops);
+ if (err)
+ return err;
}
+
+ return 0;
+}
+ALLOW_ERROR_INJECTION(xe_pt_update_ops_prepare, ERRNO);
+
+static void vma_add_fences(struct xe_vma *vma, struct dma_fence **fences,
+ int fence_count, enum dma_resv_usage usage)
+{
+ int i;
+
+ if (xe_vma_has_no_bo(vma) || xe_vma_bo(vma)->vm)
+ return;
+
+ for (i = 0; i < fence_count; ++i)
+ if (fences[i])
+ dma_resv_add_fence(xe_vma_bo(vma)->ttm.base.resv,
+ fences[i], usage);
+}
+
+static void bind_op_commit(struct xe_vm *vm, struct xe_vma *vma,
+ struct dma_fence **fences, int fence_count,
+ enum dma_resv_usage usage, u8 tile_mask,
+ bool invalidate_on_bind)
+{
+ xe_assert(vm->xe, !xe_vma_is_cpu_addr_mirror(vma));
+
+ vma_add_fences(vma, fences, fence_count, usage);
+
/* All WRITE_ONCE pair with READ_ONCE in xe_vm_has_valid_gpu_mapping() */
- WRITE_ONCE(vma->tile_present, vma->tile_present | BIT(tile->id));
+ WRITE_ONCE(vma->tile_present, vma->tile_present | tile_mask);
if (invalidate_on_bind)
WRITE_ONCE(vma->tile_invalidated,
- vma->tile_invalidated | BIT(tile->id));
+ vma->tile_invalidated | tile_mask);
else
WRITE_ONCE(vma->tile_invalidated,
- vma->tile_invalidated & ~BIT(tile->id));
- vma->tile_staged &= ~BIT(tile->id);
+ vma->tile_invalidated & ~tile_mask);
+ vma->tile_staged &= ~tile_mask;
if (xe_vma_is_userptr(vma)) {
xe_svm_assert_held_read_or_inject_write(vm);
to_userptr_vma(vma)->userptr.initial_bind = true;
@@ -2566,31 +2592,21 @@ static void bind_op_commit(struct xe_vm *vm, struct xe_tile *tile,
* Kick rebind worker if this bind triggers preempt fences and not in
* the rebind worker
*/
- if (pt_update_ops->wait_vm_bookkeep &&
+ if (usage == DMA_RESV_USAGE_KERNEL &&
xe_vm_in_preempt_fence_mode(vm) &&
!current->mm)
xe_vm_queue_rebind_worker(vm);
}
-static void unbind_op_commit(struct xe_vm *vm, struct xe_tile *tile,
- struct xe_vm_pgtable_update_ops *pt_update_ops,
- struct xe_vma *vma, struct dma_fence *fence,
- struct dma_fence *fence2)
+static void unbind_op_commit(struct xe_vm *vm, struct xe_vma *vma,
+ struct dma_fence **fences, int fence_count,
+ enum dma_resv_usage usage, u8 tile_mask)
{
- xe_tile_assert(tile, !xe_vma_is_cpu_addr_mirror(vma));
+ xe_assert(vm->xe, !xe_vma_is_cpu_addr_mirror(vma));
- if (!xe_vma_has_no_bo(vma) && !xe_vma_bo(vma)->vm) {
- dma_resv_add_fence(xe_vma_bo(vma)->ttm.base.resv, fence,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
- if (fence2)
- dma_resv_add_fence(xe_vma_bo(vma)->ttm.base.resv, fence2,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
- }
- vma->tile_present &= ~BIT(tile->id);
+ vma_add_fences(vma, fences, fence_count, usage);
+
+ vma->tile_present &= ~tile_mask;
if (!vma->tile_present) {
list_del_init(&vma->combined_links.rebind);
if (xe_vma_is_userptr(vma)) {
@@ -2605,21 +2621,19 @@ static void unbind_op_commit(struct xe_vm *vm, struct xe_tile *tile,
static void range_present_and_invalidated_tile(struct xe_vm *vm,
struct xe_svm_range *range,
- u8 tile_id)
+ u8 tile_mask)
{
/* All WRITE_ONCE pair with READ_ONCE in xe_vm_has_valid_gpu_mapping() */
lockdep_assert_held(&vm->svm.gpusvm.notifier_lock);
- WRITE_ONCE(range->tile_present, range->tile_present | BIT(tile_id));
- WRITE_ONCE(range->tile_invalidated, range->tile_invalidated & ~BIT(tile_id));
+ WRITE_ONCE(range->tile_present, range->tile_present | tile_mask);
+ WRITE_ONCE(range->tile_invalidated, range->tile_invalidated & ~tile_mask);
}
-static void op_commit(struct xe_vm *vm,
- struct xe_tile *tile,
- struct xe_vm_pgtable_update_ops *pt_update_ops,
- struct xe_vma_op *op, struct dma_fence *fence,
- struct dma_fence *fence2)
+static void op_commit(struct xe_vm *vm, struct xe_vma_op *op,
+ struct dma_fence **fences, int fence_count,
+ enum dma_resv_usage usage, u8 tile_mask)
{
xe_vm_assert_held(vm);
@@ -2629,8 +2643,8 @@ static void op_commit(struct xe_vm *vm,
(op->map.vma_flags & XE_VMA_SYSTEM_ALLOCATOR))
break;
- bind_op_commit(vm, tile, pt_update_ops, op->map.vma, fence,
- fence2, op->map.invalidate_on_bind);
+ bind_op_commit(vm, op->map.vma, fences, fence_count, usage,
+ tile_mask, op->map.invalidate_on_bind);
break;
case DRM_GPUVA_OP_REMAP:
{
@@ -2639,14 +2653,15 @@ static void op_commit(struct xe_vm *vm,
if (xe_vma_is_cpu_addr_mirror(old))
break;
- unbind_op_commit(vm, tile, pt_update_ops, old, fence, fence2);
+ unbind_op_commit(vm, old, fences, fence_count, usage,
+ tile_mask);
if (op->remap.prev && !op->remap.skip_prev)
- bind_op_commit(vm, tile, pt_update_ops, op->remap.prev,
- fence, fence2, false);
+ bind_op_commit(vm, op->remap.prev, fences, fence_count,
+ usage, tile_mask, false);
if (op->remap.next && !op->remap.skip_next)
- bind_op_commit(vm, tile, pt_update_ops, op->remap.next,
- fence, fence2, false);
+ bind_op_commit(vm, op->remap.next, fences, fence_count,
+ usage, tile_mask, false);
break;
}
case DRM_GPUVA_OP_UNMAP:
@@ -2654,8 +2669,8 @@ static void op_commit(struct xe_vm *vm,
struct xe_vma *vma = gpuva_to_vma(op->base.unmap.va);
if (!xe_vma_is_cpu_addr_mirror(vma))
- unbind_op_commit(vm, tile, pt_update_ops, vma, fence,
- fence2);
+ unbind_op_commit(vm, vma, fences, fence_count,
+ usage, tile_mask);
break;
}
case DRM_GPUVA_OP_PREFETCH:
@@ -2667,10 +2682,11 @@ static void op_commit(struct xe_vm *vm,
unsigned long i;
xa_for_each(&op->prefetch_range.range, i, range)
- range_present_and_invalidated_tile(vm, range, tile->id);
+ range_present_and_invalidated_tile(vm, range,
+ tile_mask);
} else {
- bind_op_commit(vm, tile, pt_update_ops, vma, fence,
- fence2, false);
+ bind_op_commit(vm, vma, fences, fence_count, usage,
+ tile_mask, false);
}
break;
}
@@ -2678,11 +2694,12 @@ static void op_commit(struct xe_vm *vm,
{
/* WRITE_ONCE pairs with READ_ONCE in xe_vm_has_valid_gpu_mapping() */
if (op->subop == XE_VMA_SUBOP_MAP_RANGE)
- range_present_and_invalidated_tile(vm, op->map_range.range, tile->id);
+ range_present_and_invalidated_tile(vm, op->map_range.range,
+ tile_mask);
else if (op->subop == XE_VMA_SUBOP_UNMAP_RANGE)
WRITE_ONCE(op->unmap_range.range->tile_present,
op->unmap_range.range->tile_present &
- ~BIT(tile->id));
+ ~tile_mask);
break;
}
@@ -2691,39 +2708,25 @@ static void op_commit(struct xe_vm *vm,
}
}
-static const struct xe_migrate_pt_update_ops migrate_ops = {
+static const struct xe_cpu_bind_pt_update_ops cpu_bind_ops = {
.populate = xe_vm_populate_pgtable,
- .clear = xe_migrate_clear_pgtable_callback,
+ .clear = xe_pt_clear_pgtable_callback,
.pre_commit = xe_pt_pre_commit,
};
#if IS_ENABLED(CONFIG_DRM_GPUSVM)
-static const struct xe_migrate_pt_update_ops svm_userptr_migrate_ops = {
+static const struct xe_cpu_bind_pt_update_ops svm_userptr_cpu_bind_ops = {
.populate = xe_vm_populate_pgtable,
- .clear = xe_migrate_clear_pgtable_callback,
+ .clear = xe_pt_clear_pgtable_callback,
.pre_commit = xe_pt_svm_userptr_pre_commit,
};
#else
-static const struct xe_migrate_pt_update_ops svm_userptr_migrate_ops;
+static const struct xe_cpu_bind_pt_update_ops svm_userptr_cpu_bind_ops;
#endif
-static struct xe_dep_scheduler *to_dep_scheduler(struct xe_exec_queue *q,
- struct xe_gt *gt,
- unsigned int *type)
-{
- int tile_ofs = gt_to_tile(gt)->id * (XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT + 1);
-
- if (xe_gt_is_media_type(gt))
- *type = tile_ofs + XE_EXEC_QUEUE_TLB_INVAL_MEDIA_GT;
- else
- *type = tile_ofs + XE_EXEC_QUEUE_TLB_INVAL_PRIMARY_GT;
-
- return q->tlb_inval[*type].dep_scheduler;
-}
-
/**
* xe_pt_update_ops_run() - Run PT update operations
- * @tile: Tile of PT update operations
+ * @xe: xe device.
* @vops: VMA operationa
*
* Run PT update operations which includes committing internal PT state changes,
@@ -2733,82 +2736,83 @@ static struct xe_dep_scheduler *to_dep_scheduler(struct xe_exec_queue *q,
* Return: fence on success, negative ERR_PTR on error.
*/
struct dma_fence *
-xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
+xe_pt_update_ops_run(struct xe_device *xe, struct xe_vma_ops *vops)
{
struct xe_vm *vm = vops->vm;
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &vops->pt_update_ops[tile->id];
- struct xe_exec_queue *q = pt_update_ops->q;
- struct dma_fence *fence, *ifence = NULL, *mfence = NULL;
- struct xe_tlb_inval_job *ijob = NULL, *mjob = NULL;
+ struct xe_exec_queue *q = vops->q;
+ struct dma_fence *fence;
+ struct dma_fence *ifences[XE_CPU_BIND_INVAL_JOB_COUNT] = {};
struct xe_range_fence *rfence;
+ enum dma_resv_usage usage = DMA_RESV_USAGE_BOOKKEEP;
struct xe_vma_op *op;
- unsigned int type;
- int err = 0, i;
- struct xe_migrate_pt_update update = {
- .ops = pt_update_ops->needs_svm_lock ?
- &svm_userptr_migrate_ops :
- &migrate_ops,
+ struct xe_tile *tile;
+ int err = 0, total_ops = 0, i, j;
+ u8 tile_mask = 0;
+ bool needs_invalidation = vops->flags &
+ XE_VMA_OPS_FLAG_NEEDS_INVALIDATION;
+ bool needs_svm_lock = vops->flags &
+ XE_VMA_OPS_FLAG_NEEDS_SVM_LOCK;
+ struct xe_cpu_bind_pt_update update = {
+ .ops = needs_svm_lock ? &svm_userptr_cpu_bind_ops :
+ &cpu_bind_ops,
.vops = vops,
- .tile_id = tile->id,
};
lockdep_assert_held(&vm->lock);
xe_vm_assert_held(vm);
- if (!get_current_op(pt_update_ops)) {
- xe_tile_assert(tile, xe_vm_in_fault_mode(vm));
+ for_each_tile(tile, xe, j) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &vops->pt_update_ops[j];
+ total_ops += get_current_op(pt_update_ops);
+ }
+ if (!total_ops) {
+ xe_assert(xe, xe_vm_in_fault_mode(vm));
return dma_fence_get_stub();
}
#ifdef TEST_VM_OPS_ERROR
if (vops->inject_error &&
- vm->xe->vm_inject_error_position == FORCE_OP_ERROR_RUN)
+ xe->vm_inject_error_position == FORCE_OP_ERROR_RUN)
return ERR_PTR(-ENOSPC);
#endif
- if (pt_update_ops->needs_invalidation) {
- struct xe_dep_scheduler *dep_scheduler =
- to_dep_scheduler(q, tile->primary_gt, &type);
-
- ijob = xe_tlb_inval_job_create(q, &tile->primary_gt->tlb_inval,
- dep_scheduler, vm,
- pt_update_ops->start,
- pt_update_ops->last,
- type);
- if (IS_ERR(ijob)) {
- err = PTR_ERR(ijob);
- goto kill_vm_tile1;
- }
- update.ijob = ijob;
- /*
- * Only add page reclaim for the primary GT. Media GT does not have
- * any PPC to flush, so enabling the PPC flush bit for media is
- * effectively a NOP and provides no performance benefit nor
- * interfere with primary GT.
- */
- if (xe_page_reclaim_list_valid(&pt_update_ops->prl)) {
- xe_tlb_inval_job_add_page_reclaim(ijob, &pt_update_ops->prl);
- /* Release ref from alloc, job will now handle it */
- xe_page_reclaim_list_invalidate(&pt_update_ops->prl);
- }
-
- if (tile->media_gt) {
- dep_scheduler = to_dep_scheduler(q, tile->media_gt,
- &type);
-
- mjob = xe_tlb_inval_job_create(q,
- &tile->media_gt->tlb_inval,
- dep_scheduler, vm,
- pt_update_ops->start,
- pt_update_ops->last,
- type);
- if (IS_ERR(mjob)) {
- err = PTR_ERR(mjob);
+ if (needs_invalidation) {
+ for_each_tlb_inval(q, i) {
+ struct xe_dep_scheduler *dep_scheduler =
+ q->tlb_inval[i].dep_scheduler;
+ struct xe_tile *tile =
+ &xe->tiles[i / XE_MAX_GT_PER_TILE];
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &vops->pt_update_ops[tile->id];
+ struct xe_page_reclaim_list *prl = &pt_update_ops->prl;
+ struct xe_tlb_inval_job *ijob;
+ struct xe_gt *gt = i % XE_MAX_GT_PER_TILE ?
+ tile->media_gt : tile->primary_gt;
+
+ ijob = xe_tlb_inval_job_create(q, >->tlb_inval,
+ dep_scheduler,
+ vm, vops->start,
+ vops->last, i);
+ if (IS_ERR(ijob)) {
+ err = PTR_ERR(ijob);
goto free_ijob;
}
- update.mjob = mjob;
+
+ update.ijobs[i] = ijob;
+
+ /*
+ * Only add page reclaim for the primary GT. Media GT
+ * does not have any PPC to flush, so enabling the PPC
+ * flush bit for media is effectively a NOP and provides
+ * no performance benefit nor interfere with primary GT.
+ */
+ if (xe_page_reclaim_list_valid(prl)) {
+ xe_tlb_inval_job_add_page_reclaim(ijob, prl);
+ /* Release ref from alloc, job will now handle it */
+ xe_page_reclaim_list_invalidate(prl);
+ }
}
}
@@ -2818,67 +2822,61 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
goto free_ijob;
}
- fence = xe_migrate_update_pgtables(tile->migrate, &update);
+ fence = xe_cpu_bind_update_pgtables(xe->cpu_bind, &update);
if (IS_ERR(fence)) {
err = PTR_ERR(fence);
goto free_rfence;
}
/* Point of no return - VM killed if failure after this */
- for (i = 0; i < get_current_op(pt_update_ops); ++i) {
- struct xe_vm_pgtable_update_op *pt_op =
- to_pt_op(pt_update_ops, i);
-
- xe_pt_commit(pt_op->vma, pt_op->entries,
- pt_op->num_entries,
- &pt_update_ops->pt_job_ops->deferred);
- pt_op->vma = NULL; /* skip in xe_pt_update_ops_abort */
+ for_each_tile(tile, xe, j) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &vops->pt_update_ops[j];
+
+ for (i = 0; i < get_current_op(pt_update_ops); ++i) {
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_pt_op(pt_update_ops, i);
+
+ xe_pt_commit(pt_op->vma, pt_op->entries,
+ pt_op->num_entries,
+ &pt_update_ops->pt_job_ops->deferred);
+ pt_op->vma = NULL; /* skip in xe_pt_update_ops_abort */
+ tile_mask |= BIT(tile->id);
+ }
}
- if (xe_range_fence_insert(&vm->rftree[tile->id], rfence,
+ if (xe_range_fence_insert(&vm->rftree, rfence,
&xe_range_fence_kfree_ops,
- pt_update_ops->start,
- pt_update_ops->last, fence))
+ vops->start, vops->last, fence))
dma_fence_wait(fence, false);
- if (ijob)
- ifence = xe_tlb_inval_job_push(ijob, tile->migrate, fence);
- if (mjob)
- mfence = xe_tlb_inval_job_push(mjob, tile->migrate, fence);
+ if (vops->flags & XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP)
+ usage = DMA_RESV_USAGE_KERNEL;
- if (!mjob && !ijob) {
- dma_resv_add_fence(xe_vm_resv(vm), fence,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
-
- list_for_each_entry(op, &vops->list, link)
- op_commit(vops->vm, tile, pt_update_ops, op, fence, NULL);
- } else if (ijob && !mjob) {
- dma_resv_add_fence(xe_vm_resv(vm), ifence,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
+ if (!needs_invalidation) {
+ dma_resv_add_fence(xe_vm_resv(vm), fence, usage);
list_for_each_entry(op, &vops->list, link)
- op_commit(vops->vm, tile, pt_update_ops, op, ifence, NULL);
+ op_commit(vops->vm, op, &fence, 1, usage, tile_mask);
} else {
- dma_resv_add_fence(xe_vm_resv(vm), ifence,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
+ for (i = 0; i < XE_CPU_BIND_INVAL_JOB_COUNT; ++i) {
+ if (!update.ijobs[i])
+ continue;
+
+ ifences[i] = xe_tlb_inval_job_push(update.ijobs[i],
+ fence);
+ xe_assert(xe, !IS_ERR_OR_NULL(ifences[i]));
- dma_resv_add_fence(xe_vm_resv(vm), mfence,
- pt_update_ops->wait_vm_bookkeep ?
- DMA_RESV_USAGE_KERNEL :
- DMA_RESV_USAGE_BOOKKEEP);
+ dma_resv_add_fence(xe_vm_resv(vm), ifences[i], usage);
+ }
list_for_each_entry(op, &vops->list, link)
- op_commit(vops->vm, tile, pt_update_ops, op, ifence,
- mfence);
+ op_commit(vops->vm, op, ifences,
+ XE_CPU_BIND_INVAL_JOB_COUNT, usage,
+ tile_mask);
}
- if (pt_update_ops->needs_svm_lock)
+ if (needs_svm_lock)
xe_pt_svm_userptr_notifier_unlock(vm);
/*
@@ -2888,21 +2886,18 @@ xe_pt_update_ops_run(struct xe_tile *tile, struct xe_vma_ops *vops)
if (!(q->flags & EXEC_QUEUE_FLAG_MIGRATE))
xe_exec_queue_last_fence_set(q, vm, fence);
- xe_tlb_inval_job_put(mjob);
- xe_tlb_inval_job_put(ijob);
- dma_fence_put(ifence);
- dma_fence_put(mfence);
+ for (i = 0; i < XE_CPU_BIND_INVAL_JOB_COUNT; ++i) {
+ xe_tlb_inval_job_put(update.ijobs[i]);
+ dma_fence_put(ifences[i]);
+ }
return fence;
free_rfence:
kfree(rfence);
free_ijob:
- xe_tlb_inval_job_put(mjob);
- xe_tlb_inval_job_put(ijob);
-kill_vm_tile1:
- if (err != -EAGAIN && err != -ENODATA && tile->id)
- xe_vm_kill(vops->vm, false);
+ for (i = 0; i < XE_CPU_BIND_INVAL_JOB_COUNT; ++i)
+ xe_tlb_inval_job_put(update.ijobs[i]);
return ERR_PTR(err);
}
@@ -2910,52 +2905,65 @@ ALLOW_ERROR_INJECTION(xe_pt_update_ops_run, ERRNO);
/**
* xe_pt_update_ops_fini() - Finish PT update operations
- * @tile: Tile of PT update operations
+ * @xe: xe device.
* @vops: VMA operations
*
* Finish PT update operations by committing to destroy page table memory
*/
-void xe_pt_update_ops_fini(struct xe_tile *tile, struct xe_vma_ops *vops)
+void xe_pt_update_ops_fini(struct xe_device *xe, struct xe_vma_ops *vops)
{
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &vops->pt_update_ops[tile->id];
+ struct xe_tile *tile;
+ int id;
+
+ for_each_tile(tile, xe, id) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &vops->pt_update_ops[id];
- xe_page_reclaim_entries_put(pt_update_ops->prl.entries);
+ if (!pt_update_ops->num_ops)
+ continue;
+
+ xe_page_reclaim_entries_put(pt_update_ops->prl.entries);
+ }
}
/**
* xe_pt_update_ops_abort() - Abort PT update operations
- * @tile: Tile of PT update operations
+ * @xe: xe device.
* @vops: VMA operationa
*
* Abort PT update operations by unwinding internal PT state
*/
-void xe_pt_update_ops_abort(struct xe_tile *tile, struct xe_vma_ops *vops)
+void xe_pt_update_ops_abort(struct xe_device *xe, struct xe_vma_ops *vops)
{
- struct xe_vm_pgtable_update_ops *pt_update_ops =
- &vops->pt_update_ops[tile->id];
- int i;
+ struct xe_tile *tile;
+ int id;
lockdep_assert_held(&vops->vm->lock);
xe_vm_assert_held(vops->vm);
- for (i = pt_update_ops->num_ops - 1; i >= 0; --i) {
- struct xe_vm_pgtable_update_op *pt_op =
- to_pt_op(pt_update_ops, i);
+ for_each_tile(tile, xe, id) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &vops->pt_update_ops[id];
+ int i;
- if (!pt_op->vma || i >= get_current_op(pt_update_ops))
- continue;
+ for (i = pt_update_ops->num_ops - 1; i >= 0; --i) {
+ struct xe_vm_pgtable_update_op *pt_op =
+ to_pt_op(pt_update_ops, i);
- if (pt_op->bind)
- xe_pt_abort_bind(pt_op->vma, pt_op->entries,
- pt_op->num_entries,
- pt_op->rebind);
- else
- xe_pt_abort_unbind(pt_op->vma, pt_op->entries,
- pt_op->num_entries);
+ if (!pt_op->vma || i >= get_current_op(pt_update_ops))
+ continue;
+
+ if (pt_op->bind)
+ xe_pt_abort_bind(pt_op->vma, pt_op->entries,
+ pt_op->num_entries,
+ pt_op->rebind);
+ else
+ xe_pt_abort_unbind(pt_op->vma, pt_op->entries,
+ pt_op->num_entries);
+ }
}
- xe_pt_update_ops_fini(tile, vops);
+ xe_pt_update_ops_fini(xe, vops);
}
/**
diff --git a/drivers/gpu/drm/xe/xe_pt.h b/drivers/gpu/drm/xe/xe_pt.h
index 5faddb8e700c..cd78141fb81c 100644
--- a/drivers/gpu/drm/xe/xe_pt.h
+++ b/drivers/gpu/drm/xe/xe_pt.h
@@ -39,11 +39,11 @@ void xe_pt_destroy(struct xe_pt *pt, u32 flags, struct llist_head *deferred);
void xe_pt_clear(struct xe_device *xe, struct xe_pt *pt);
-int xe_pt_update_ops_prepare(struct xe_tile *tile, struct xe_vma_ops *vops);
-struct dma_fence *xe_pt_update_ops_run(struct xe_tile *tile,
+int xe_pt_update_ops_prepare(struct xe_device *xe, struct xe_vma_ops *vops);
+struct dma_fence *xe_pt_update_ops_run(struct xe_device *xe,
struct xe_vma_ops *vops);
-void xe_pt_update_ops_fini(struct xe_tile *tile, struct xe_vma_ops *vops);
-void xe_pt_update_ops_abort(struct xe_tile *tile, struct xe_vma_ops *vops);
+void xe_pt_update_ops_fini(struct xe_device *xe, struct xe_vma_ops *vops);
+void xe_pt_update_ops_abort(struct xe_device *xe, struct xe_vma_ops *vops);
bool xe_pt_zap_ptes(struct xe_tile *tile, struct xe_vma *vma);
bool xe_pt_zap_ptes_range(struct xe_tile *tile, struct xe_vm *vm,
diff --git a/drivers/gpu/drm/xe/xe_pt_types.h b/drivers/gpu/drm/xe/xe_pt_types.h
index ccab6613385f..0d4bac22ee6c 100644
--- a/drivers/gpu/drm/xe/xe_pt_types.h
+++ b/drivers/gpu/drm/xe/xe_pt_types.h
@@ -120,8 +120,6 @@ struct xe_pt_job_ops {
struct xe_vm_pgtable_update_ops {
/** @pt_job_ops: PT update operations dynamic allocation*/
struct xe_pt_job_ops *pt_job_ops;
- /** @q: exec queue for PT operations */
- struct xe_exec_queue *q;
/** @prl: embedded page reclaim list */
struct xe_page_reclaim_list prl;
/** @start: start address of ops */
@@ -134,18 +132,6 @@ struct xe_vm_pgtable_update_ops {
bool needs_svm_lock;
/** @needs_invalidation: Needs invalidation */
bool needs_invalidation;
- /**
- * @wait_vm_bookkeep: PT operations need to wait until VM is idle
- * (bookkeep dma-resv slots are idle) and stage all future VM activity
- * behind these operations (install PT operations into VM kernel
- * dma-resv slot).
- */
- bool wait_vm_bookkeep;
- /**
- * @wait_vm_kernel: PT operations need to wait until VM kernel dma-resv
- * slots are idle.
- */
- bool wait_vm_kernel;
};
#endif
diff --git a/drivers/gpu/drm/xe/xe_sched_job.c b/drivers/gpu/drm/xe/xe_sched_job.c
index 20992ea816e6..59f7215f27b1 100644
--- a/drivers/gpu/drm/xe/xe_sched_job.c
+++ b/drivers/gpu/drm/xe/xe_sched_job.c
@@ -73,8 +73,9 @@ static void job_free(struct xe_sched_job *job)
struct xe_exec_queue *q = job->q;
bool is_migration = xe_sched_job_is_migration(q);
- kmem_cache_free(xe_exec_queue_is_parallel(job->q) || is_migration ?
- xe_sched_job_parallel_slab : xe_sched_job_slab, job);
+ kmem_cache_free(job->is_pt_job || xe_exec_queue_is_parallel(job->q) ||
+ is_migration ? xe_sched_job_parallel_slab :
+ xe_sched_job_slab, job);
}
static struct xe_device *job_to_xe(struct xe_sched_job *job)
@@ -127,10 +128,12 @@ struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
xe_assert(xe, batch_addr ||
q->flags & (EXEC_QUEUE_FLAG_VM | EXEC_QUEUE_FLAG_MIGRATE));
- job = job_alloc(xe_exec_queue_is_parallel(q) || is_migration);
+ job = job_alloc(!batch_addr || xe_exec_queue_is_parallel(q) ||
+ is_migration);
if (!job)
return ERR_PTR(-ENOMEM);
+ job->is_pt_job = !batch_addr;
job->q = q;
job->sample_timestamp = U64_MAX;
kref_init(&job->refcount);
@@ -143,7 +146,6 @@ struct xe_sched_job *xe_sched_job_create(struct xe_exec_queue *q,
if (!batch_addr) {
job->fence = dma_fence_get_stub();
- job->is_pt_job = true;
} else {
for (i = 0; i < q->width; ++i) {
struct dma_fence *fence = xe_lrc_alloc_seqno_fence();
diff --git a/drivers/gpu/drm/xe/xe_sched_job_types.h b/drivers/gpu/drm/xe/xe_sched_job_types.h
index 5e1824c36c74..9f527ac6df3e 100644
--- a/drivers/gpu/drm/xe/xe_sched_job_types.h
+++ b/drivers/gpu/drm/xe/xe_sched_job_types.h
@@ -14,7 +14,7 @@ struct dma_fence;
struct dma_fence_chain;
struct xe_exec_queue;
-struct xe_migrate_pt_update_ops;
+struct xe_cpu_bind_pt_update_ops;
struct xe_pt_job_ops;
struct xe_tile;
struct xe_vm;
@@ -25,12 +25,11 @@ struct xe_vm;
struct xe_pt_update_args {
/** @vm: VM which is being bound */
struct xe_vm *vm;
- /** @tile: Tile which page tables belong to */
- struct xe_tile *tile;
- /** @ops: Migrate PT update ops */
- const struct xe_migrate_pt_update_ops *ops;
+ /** @ops: CPU bind PT update ops */
+ const struct xe_cpu_bind_pt_update_ops *ops;
+#define XE_PT_UPDATE_JOB_OPS_COUNT 2
/** @pt_job_ops: PT job ops state */
- struct xe_pt_job_ops *pt_job_ops;
+ struct xe_pt_job_ops *pt_job_ops[XE_PT_UPDATE_JOB_OPS_COUNT];
};
/**
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_job.c b/drivers/gpu/drm/xe/xe_tlb_inval_job.c
index 81f560068d3c..7378cfe6e855 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_job.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_job.c
@@ -4,6 +4,7 @@
*/
#include "xe_assert.h"
+#include "xe_cpu_bind.h"
#include "xe_dep_job_types.h"
#include "xe_dep_scheduler.h"
#include "xe_exec_queue.h"
@@ -12,7 +13,6 @@
#include "xe_page_reclaim.h"
#include "xe_tlb_inval.h"
#include "xe_tlb_inval_job.h"
-#include "xe_migrate.h"
#include "xe_pm.h"
#include "xe_vm.h"
@@ -218,7 +218,6 @@ int xe_tlb_inval_job_alloc_dep(struct xe_tlb_inval_job *job)
/**
* xe_tlb_inval_job_push() - TLB invalidation job push
* @job: TLB invalidation job to push
- * @m: The migration object being used
* @fence: Dependency for TLB invalidation job
*
* Pushes a TLB invalidation job for execution, using @fence as a dependency.
@@ -230,11 +229,11 @@ int xe_tlb_inval_job_alloc_dep(struct xe_tlb_inval_job *job)
* Return: Job's finished fence on success, cannot fail
*/
struct dma_fence *xe_tlb_inval_job_push(struct xe_tlb_inval_job *job,
- struct xe_migrate *m,
struct dma_fence *fence)
{
struct xe_tlb_inval_fence *ifence =
container_of(job->fence, typeof(*ifence), base);
+ struct xe_cpu_bind *cpu_bind = gt_to_xe(job->q->gt)->cpu_bind;
if (!dma_fence_is_signaled(fence)) {
void *ptr;
@@ -258,11 +257,11 @@ struct dma_fence *xe_tlb_inval_job_push(struct xe_tlb_inval_job *job,
job->fence_armed = true;
/*
- * We need the migration lock to protect the job's seqno and the spsc
- * queue, only taken on migration queue, user queues protected dma-resv
+ * We need the cpu_bind lock to protect the job's seqno and the spsc
+ * queue, only taken on cpu_bind queue, user queues protected dma-resv
* VM lock.
*/
- xe_migrate_job_lock(m, job->q);
+ xe_cpu_bind_job_lock(cpu_bind, job->q);
/* Creation ref pairs with put in xe_tlb_inval_job_destroy */
xe_tlb_inval_fence_init(job->tlb_inval, ifence, false);
@@ -281,7 +280,7 @@ struct dma_fence *xe_tlb_inval_job_push(struct xe_tlb_inval_job *job,
&job->dep.drm.s_fence->finished,
job->idx);
- xe_migrate_job_unlock(m, job->q);
+ xe_cpu_bind_job_unlock(cpu_bind, job->q);
/*
* Not using job->fence, as it has its own dma-fence context, which does
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_job.h b/drivers/gpu/drm/xe/xe_tlb_inval_job.h
index 2a4478f529e6..97e032ea21c3 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_job.h
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_job.h
@@ -11,7 +11,6 @@
struct dma_fence;
struct xe_dep_scheduler;
struct xe_exec_queue;
-struct xe_migrate;
struct xe_page_reclaim_list;
struct xe_tlb_inval;
struct xe_tlb_inval_job;
@@ -28,7 +27,6 @@ void xe_tlb_inval_job_add_page_reclaim(struct xe_tlb_inval_job *job,
int xe_tlb_inval_job_alloc_dep(struct xe_tlb_inval_job *job);
struct dma_fence *xe_tlb_inval_job_push(struct xe_tlb_inval_job *job,
- struct xe_migrate *m,
struct dma_fence *fence);
void xe_tlb_inval_job_get(struct xe_tlb_inval_job *job);
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index feb67a565fb0..0f542f47b9a8 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -24,6 +24,7 @@
#include "regs/xe_gtt_defs.h"
#include "xe_assert.h"
#include "xe_bo.h"
+#include "xe_cpu_bind.h"
#include "xe_device.h"
#include "xe_drm_client.h"
#include "xe_exec_queue.h"
@@ -796,8 +797,6 @@ int xe_vm_rebind(struct xe_vm *vm, bool rebind_worker)
struct xe_vma *vma, *next;
struct xe_vma_ops vops;
struct xe_vma_op *op, *next_op;
- struct xe_tile *tile;
- u8 id;
int err;
lockdep_assert_held(&vm->lock);
@@ -805,12 +804,9 @@ int xe_vm_rebind(struct xe_vm *vm, bool rebind_worker)
list_empty(&vm->rebind_list))
return 0;
- xe_vma_ops_init(&vops, vm, NULL, NULL, 0);
- for_each_tile(tile, vm->xe, id) {
- vops.pt_update_ops[id].wait_vm_bookkeep = true;
- vops.pt_update_ops[id].q =
- xe_migrate_bind_queue(tile->migrate);
- }
+ xe_vma_ops_init(&vops, vm, xe_cpu_bind_queue(vm->xe->cpu_bind),
+ NULL, 0);
+ vops.flags |= XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
xe_vm_assert_held(vm);
list_for_each_entry(vma, &vm->rebind_list, combined_links.rebind) {
@@ -855,21 +851,16 @@ struct dma_fence *xe_vma_rebind(struct xe_vm *vm, struct xe_vma *vma, u8 tile_ma
struct dma_fence *fence = NULL;
struct xe_vma_ops vops;
struct xe_vma_op *op, *next_op;
- struct xe_tile *tile;
- u8 id;
int err;
lockdep_assert_held(&vm->lock);
xe_vm_assert_held(vm);
xe_assert(vm->xe, xe_vm_in_fault_mode(vm));
- xe_vma_ops_init(&vops, vm, NULL, NULL, 0);
- vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT;
- for_each_tile(tile, vm->xe, id) {
- vops.pt_update_ops[id].wait_vm_bookkeep = true;
- vops.pt_update_ops[tile->id].q =
- xe_migrate_bind_queue(tile->migrate);
- }
+ xe_vma_ops_init(&vops, vm, xe_cpu_bind_queue(vm->xe->cpu_bind),
+ NULL, 0);
+ vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT |
+ XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
err = xe_vm_ops_add_rebind(&vops, vma, tile_mask);
if (err)
@@ -945,8 +936,6 @@ struct dma_fence *xe_vm_range_rebind(struct xe_vm *vm,
struct dma_fence *fence = NULL;
struct xe_vma_ops vops;
struct xe_vma_op *op, *next_op;
- struct xe_tile *tile;
- u8 id;
int err;
lockdep_assert_held(&range->lock);
@@ -955,13 +944,10 @@ struct dma_fence *xe_vm_range_rebind(struct xe_vm *vm,
xe_assert(vm->xe, xe_vm_in_fault_mode(vm));
xe_assert(vm->xe, xe_vma_is_cpu_addr_mirror(vma));
- xe_vma_ops_init(&vops, vm, NULL, NULL, 0);
- vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT;
- for_each_tile(tile, vm->xe, id) {
- vops.pt_update_ops[id].wait_vm_bookkeep = true;
- vops.pt_update_ops[tile->id].q =
- xe_migrate_bind_queue(tile->migrate);
- }
+ xe_vma_ops_init(&vops, vm, xe_cpu_bind_queue(vm->xe->cpu_bind),
+ NULL, 0);
+ vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT |
+ XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
err = xe_vm_ops_add_range_rebind(&vops, vma, range, tile_mask);
if (err)
@@ -1028,8 +1014,6 @@ struct dma_fence *xe_vm_range_unbind(struct xe_vm *vm,
struct dma_fence *fence = NULL;
struct xe_vma_ops vops;
struct xe_vma_op *op, *next_op;
- struct xe_tile *tile;
- u8 id;
int err;
lockdep_assert_held(&range->lock);
@@ -1040,12 +1024,9 @@ struct dma_fence *xe_vm_range_unbind(struct xe_vm *vm,
if (!range->tile_present)
return dma_fence_get_stub();
- xe_vma_ops_init(&vops, vm, NULL, NULL, 0);
- for_each_tile(tile, vm->xe, id) {
- vops.pt_update_ops[id].wait_vm_bookkeep = true;
- vops.pt_update_ops[tile->id].q =
- xe_migrate_bind_queue(tile->migrate);
- }
+ xe_vma_ops_init(&vops, vm, xe_cpu_bind_queue(vm->xe->cpu_bind),
+ NULL, 0);
+ vops.flags |= XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
err = xe_vm_ops_add_range_unbind(&vops, range);
if (err)
@@ -1716,9 +1697,7 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
init_rwsem(&vm->exec_queues.lock);
xe_vm_init_prove_locking(xe, vm);
-
- for_each_tile(tile, xe, id)
- xe_range_fence_tree_init(&vm->rftree[id]);
+ xe_range_fence_tree_init(&vm->rftree);
vm->pt_ops = &xelp_pt_ops;
@@ -1860,8 +1839,7 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
xe_svm_fini(vm);
err_no_resv:
mutex_destroy(&vm->snap_mutex);
- for_each_tile(tile, xe, id)
- xe_range_fence_tree_fini(&vm->rftree[id]);
+ xe_range_fence_tree_fini(&vm->rftree);
ttm_lru_bulk_move_fini(&xe->ttm, &vm->lru_bulk_move);
if (vm->xef)
xe_file_put(vm->xef);
@@ -1921,10 +1899,8 @@ void xe_vm_close_and_put(struct xe_vm *vm)
{
LIST_HEAD(contested);
struct xe_device *xe = vm->xe;
- struct xe_tile *tile;
struct xe_vma *vma, *next_vma;
struct drm_gpuva *gpuva, *next;
- u8 id;
xe_assert(xe, !vm->preempt.num_exec_queues);
@@ -2013,8 +1989,7 @@ void xe_vm_close_and_put(struct xe_vm *vm)
xe_vm_clear_fault_entries(vm);
- for_each_tile(tile, xe, id)
- xe_range_fence_tree_fini(&vm->rftree[id]);
+ xe_range_fence_tree_fini(&vm->rftree);
xe_vm_put(vm);
}
@@ -3509,25 +3484,18 @@ static void trace_xe_vm_ops_execute(struct xe_vma_ops *vops)
op_trace(op);
}
-static int vm_ops_setup_tile_args(struct xe_vm *vm, struct xe_vma_ops *vops)
+static int vm_ops_setup(struct xe_vm *vm, struct xe_vma_ops *vops)
{
- struct xe_exec_queue *q = vops->q;
struct xe_tile *tile;
int number_tiles = 0;
u8 id;
- for_each_tile(tile, vm->xe, id) {
+ for_each_tile(tile, vm->xe, id)
if (vops->pt_update_ops[id].num_ops)
++number_tiles;
- if (vops->pt_update_ops[id].q)
- continue;
-
- if (q)
- vops->pt_update_ops[id].q = q;
- else
- vops->pt_update_ops[id].q = vm->q;
- }
+ if (!vops->q)
+ vops->q = vm->q;
return number_tiles;
}
@@ -3535,22 +3503,17 @@ static int vm_ops_setup_tile_args(struct xe_vm *vm, struct xe_vma_ops *vops)
static struct dma_fence *ops_execute(struct xe_vm *vm,
struct xe_vma_ops *vops)
{
- struct xe_tile *tile;
+ struct xe_device *xe = vm->xe;
struct dma_fence *fence = NULL;
struct dma_fence **fences = NULL;
struct dma_fence_array *cf = NULL;
- int number_tiles = 0, current_fence = 0, n_fence = 0, err, i;
- u8 id;
+ int current_fence = 0, n_fence = 1, err, i;
- number_tiles = vm_ops_setup_tile_args(vm, vops);
- if (number_tiles == 0)
+ if (!vm_ops_setup(vm, vops))
return ERR_PTR(-ENODATA);
- for_each_tile(tile, vm->xe, id)
- ++n_fence;
-
if (!(vops->flags & XE_VMA_OPS_FLAG_SKIP_TLB_WAIT)) {
- for_each_tlb_inval(vops->pt_update_ops[0].q, i)
+ for_each_tlb_inval(vops->q, i)
++n_fence;
}
@@ -3566,69 +3529,39 @@ static struct dma_fence *ops_execute(struct xe_vm *vm,
goto err_out;
}
- for_each_tile(tile, vm->xe, id) {
- if (!vops->pt_update_ops[id].num_ops)
- continue;
-
- err = xe_pt_update_ops_prepare(tile, vops);
- if (err) {
- fence = ERR_PTR(err);
- goto err_out;
- }
+ err = xe_pt_update_ops_prepare(xe, vops);
+ if (err) {
+ fence = ERR_PTR(err);
+ goto err_out;
}
trace_xe_vm_ops_execute(vops);
- for_each_tile(tile, vm->xe, id) {
- struct xe_exec_queue *q = vops->pt_update_ops[tile->id].q;
-
- fence = NULL;
- if (!vops->pt_update_ops[id].num_ops)
- goto collect_fences;
-
- fence = xe_pt_update_ops_run(tile, vops);
- if (IS_ERR(fence))
- goto err_out;
-
-collect_fences:
- fences[current_fence++] = fence ?: dma_fence_get_stub();
- if (vops->flags & XE_VMA_OPS_FLAG_SKIP_TLB_WAIT)
- continue;
-
- xe_migrate_job_lock(tile->migrate, q);
- for_each_tlb_inval(q, i) {
- if (i >= (tile->id + 1) * XE_MAX_GT_PER_TILE ||
- i < tile->id * XE_MAX_GT_PER_TILE)
- continue;
+ fence = xe_pt_update_ops_run(xe, vops);
+ if (IS_ERR(fence))
+ goto err_out;
+ fences[current_fence++] = fence;
- fences[current_fence++] = fence ?
- xe_exec_queue_tlb_inval_last_fence_get(q, vm, i) :
- dma_fence_get_stub();
- }
- xe_migrate_job_unlock(tile->migrate, q);
+ if (!(vops->flags & XE_VMA_OPS_FLAG_SKIP_TLB_WAIT)) {
+ xe_cpu_bind_job_lock(xe->cpu_bind, vops->q);
+ for_each_tlb_inval(vops->q, i)
+ fences[current_fence++] =
+ xe_exec_queue_tlb_inval_last_fence_get(vops->q,
+ vm, i);
+ xe_cpu_bind_job_unlock(xe->cpu_bind, vops->q);
}
- xe_assert(vm->xe, current_fence == n_fence);
+ xe_assert(xe, current_fence == n_fence);
dma_fence_array_init(cf, n_fence, fences, dma_fence_context_alloc(1),
1);
fence = &cf->base;
- for_each_tile(tile, vm->xe, id) {
- if (!vops->pt_update_ops[id].num_ops)
- continue;
-
- xe_pt_update_ops_fini(tile, vops);
- }
+ xe_pt_update_ops_fini(xe, vops);
return fence;
err_out:
- for_each_tile(tile, vm->xe, id) {
- if (!vops->pt_update_ops[id].num_ops)
- continue;
-
- xe_pt_update_ops_abort(tile, vops);
- }
+ xe_pt_update_ops_abort(xe, vops);
while (current_fence)
dma_fence_put(fences[--current_fence]);
kfree(fences);
@@ -3940,6 +3873,8 @@ static void xe_vma_ops_init(struct xe_vma_ops *vops, struct xe_vm *vm,
vops->syncs = syncs;
vops->num_syncs = num_syncs;
vops->flags = 0;
+ vops->start = ~0x0ull;
+ vops->last = 0x0ull;
}
static int xe_vm_bind_ioctl_validate_bo(struct xe_device *xe, struct xe_bo *bo,
diff --git a/drivers/gpu/drm/xe/xe_vm_types.h b/drivers/gpu/drm/xe/xe_vm_types.h
index c91cb13fc4f1..fb9305124679 100644
--- a/drivers/gpu/drm/xe/xe_vm_types.h
+++ b/drivers/gpu/drm/xe/xe_vm_types.h
@@ -317,7 +317,7 @@ struct xe_vm {
* @rftree: range fence tree to track updates to page table structure.
* Used to implement conflict tracking between independent bind engines.
*/
- struct xe_range_fence_tree rftree[XE_MAX_TILES_PER_DEVICE];
+ struct xe_range_fence_tree rftree;
const struct xe_pt_ops *pt_ops;
@@ -557,6 +557,10 @@ struct xe_vma_ops {
u32 num_syncs;
/** @pt_update_ops: page table update operations */
struct xe_vm_pgtable_update_ops pt_update_ops[XE_MAX_TILES_PER_DEVICE];
+ /** @start: start address of ops */
+ u64 start;
+ /** @last: last address of ops */
+ u64 last;
/** @flag: signify the properties within xe_vma_ops*/
#define XE_VMA_OPS_FLAG_HAS_SVM_PREFETCH BIT(0)
#define XE_VMA_OPS_FLAG_MADVISE BIT(1)
@@ -566,6 +570,10 @@ struct xe_vma_ops {
#define XE_VMA_OPS_FLAG_MODIFIES_GPUVA BIT(5)
#define XE_VMA_OPS_FLAG_DOWNGRADE_LOCK BIT(6)
#define XE_VMA_OPS_FLAG_HAS_SVM_VALID_RANGE BIT(7)
+#define XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP BIT(8)
+#define XE_VMA_OPS_FLAG_WAIT_VM_KERNEL BIT(9)
+#define XE_VMA_OPS_FLAG_NEEDS_INVALIDATION BIT(10)
+#define XE_VMA_OPS_FLAG_NEEDS_SVM_LOCK BIT(11)
u32 flags;
#ifdef TEST_VM_OPS_ERROR
/** @inject_error: inject error to test error handling */
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (15 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 16/24] drm/xe: Add CPU bind layer Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 6:25 ` sashiko-bot
2026-09-25 9:51 ` Francois Dugast
2026-09-25 4:53 ` [PATCH v7 18/24] drm/xe: Add ULLS support to LRC Matthew Brost
` (11 subsequent siblings)
28 siblings, 2 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe
Some multi-tile devices may want to mirror page tables across tiles for
memory-bandwidth reasons, while others may not. Add a device flag that
allows enabling or disabling page-table mirroring across tiles.
Setting the flag to true (the existing behavior) on PVC, but both modes
have been tested and are working on PVC.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
---
v7:
- Rework pt_mirroring_disabled_for_tile (Himal)
---
drivers/gpu/drm/xe/xe_device_types.h | 2 ++
drivers/gpu/drm/xe/xe_migrate.c | 5 ++--
drivers/gpu/drm/xe/xe_pci.c | 2 ++
drivers/gpu/drm/xe/xe_pci_types.h | 3 ++-
drivers/gpu/drm/xe/xe_pt.c | 36 +++++++++++++++++++++++++--
drivers/gpu/drm/xe/xe_vm.c | 37 +++++++++++++++++++++++++---
drivers/gpu/drm/xe/xe_vm.h | 3 +++
7 files changed, 79 insertions(+), 9 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_device_types.h b/drivers/gpu/drm/xe/xe_device_types.h
index aff2e5c8b656..cd9ed8ff2940 100644
--- a/drivers/gpu/drm/xe/xe_device_types.h
+++ b/drivers/gpu/drm/xe/xe_device_types.h
@@ -229,6 +229,8 @@ struct xe_device {
u8 has_usm:1;
/** @info.has_64bit_timestamp: Device supports 64-bit timestamps */
u8 has_64bit_timestamp:1;
+ /** @info.has_pt_mirror: Device has PT mirroring across tiles */
+ u8 has_pt_mirror:1;
/** @info.is_dgfx: is discrete device */
u8 is_dgfx:1;
/** @info.needs_scratch: needs scratch page for oob prefetch to work */
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index f7e1a81434b2..471ae5741836 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -246,7 +246,8 @@ static void xe_migrate_prepare_vm(struct xe_tile *tile, struct xe_migrate *m,
struct xe_device *xe = tile_to_xe(tile);
u16 pat_index = xe_cache_pat_idx(xe, XE_CACHE_WB);
u8 id = tile->id;
- u32 num_entries = NUM_PT_SLOTS, num_level = vm->pt_root[id]->level;
+ u32 num_entries = NUM_PT_SLOTS, num_level =
+ xe_vm_pt_root(vm, id)->level;
#define VRAM_IDENTITY_MAP_PT_COUNT 4
u32 num_setup = num_level + VRAM_IDENTITY_MAP_PT_COUNT;
#undef VRAM_IDENTITY_MAP_PT_COUNT
@@ -258,7 +259,7 @@ static void xe_migrate_prepare_vm(struct xe_tile *tile, struct xe_migrate *m,
u64 l1_pt_ofs = xe_bo_size(bo) - 5 * XE_PAGE_SIZE;
entry = vm->pt_ops->pde_encode_bo(bo, l1_pt_ofs);
- xe_pt_write(xe, &vm->pt_root[id]->bo->vmap, 0, entry);
+ xe_pt_write(xe, &xe_vm_pt_root(vm, id)->bo->vmap, 0, entry);
map_ofs = (num_entries - num_setup) * XE_PAGE_SIZE;
diff --git a/drivers/gpu/drm/xe/xe_pci.c b/drivers/gpu/drm/xe/xe_pci.c
index d64e873d9e71..28f37e034e02 100644
--- a/drivers/gpu/drm/xe/xe_pci.c
+++ b/drivers/gpu/drm/xe/xe_pci.c
@@ -371,6 +371,7 @@ static const __maybe_unused struct xe_device_desc pvc_desc = {
.has_display = false,
.has_drm_ras = true,
.has_gsc_nvm = 1,
+ .has_pt_mirror = 1,
.has_heci_gscfi = 1,
.max_gt_per_tile = 1,
.max_remote_tiles = 1,
@@ -807,6 +808,7 @@ static int xe_info_init_early(struct xe_device *xe,
xe->info.has_mert = desc->has_mert;
xe->info.has_page_reclaim_hw_assist = desc->has_page_reclaim_hw_assist;
xe->info.has_pre_prod_wa = desc->has_pre_prod_wa;
+ xe->info.has_pt_mirror = desc->has_pt_mirror;
xe->info.has_pxp = desc->has_pxp;
xe->info.has_soc_remapper_sysctrl = desc->has_soc_remapper_sysctrl;
xe->info.has_soc_remapper_telem = desc->has_soc_remapper_telem;
diff --git a/drivers/gpu/drm/xe/xe_pci_types.h b/drivers/gpu/drm/xe/xe_pci_types.h
index fed509ff601e..71068cdb3558 100644
--- a/drivers/gpu/drm/xe/xe_pci_types.h
+++ b/drivers/gpu/drm/xe/xe_pci_types.h
@@ -52,8 +52,9 @@ struct xe_device_desc {
u8 has_mbx_power_limits:1;
u8 has_mbx_thermal_info:1;
u8 has_mert:1;
- u8 has_pre_prod_wa:1;
u8 has_page_reclaim_hw_assist:1;
+ u8 has_pre_prod_wa:1;
+ u8 has_pt_mirror:1;
u8 has_pxp:1;
u8 has_soc_remapper_sysctrl:1;
u8 has_soc_remapper_telem:1;
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index fbefcfc52dca..dcef29b56d63 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -804,7 +804,7 @@ xe_pt_stage_bind(struct xe_tile *tile, struct xe_vma *vma,
.wupd.entries = entries,
.clear_pt = clear_pt,
};
- struct xe_pt *pt = vm->pt_root[tile->id];
+ struct xe_pt *pt = xe_vm_pt_root(vm, tile->id);
int ret;
bool is_purged = false;
@@ -1005,6 +1005,11 @@ static int xe_pt_zap_ptes_entry(struct xe_ptw *parent, pgoff_t offset,
return 0;
}
+static bool pt_mirroring_disabled_for_tile(struct xe_vm *vm, u8 tile_id)
+{
+ return !vm->xe->info.has_pt_mirror && tile_id;
+}
+
static const struct xe_pt_walk_ops xe_pt_zap_ptes_ops = {
.pt_entry = xe_pt_zap_ptes_entry,
};
@@ -1046,6 +1051,9 @@ bool xe_pt_zap_ptes(struct xe_tile *tile, struct xe_vma *vma)
if (!(pt_mask & BIT(tile->id)))
return false;
+ if (pt_mirroring_disabled_for_tile(xe_vma_vm(vma), tile->id))
+ return true;
+
(void)xe_pt_walk_shared(&pt->base, pt->level, xe_vma_start(vma),
xe_vma_end(vma), &xe_walk.base);
@@ -1098,6 +1106,9 @@ bool xe_pt_zap_ptes_range(struct xe_tile *tile, struct xe_vm *vm,
if (!(pt_mask & BIT(tile->id)))
return false;
+ if (pt_mirroring_disabled_for_tile(vm, tile->id))
+ return true;
+
(void)xe_pt_walk_shared(&pt->base, pt->level, xe_svm_range_start(range),
xe_svm_range_end(range), &xe_walk.base);
@@ -1987,7 +1998,7 @@ static unsigned int xe_pt_stage_unbind(struct xe_tile *tile,
.wupd.entries = entries,
.prl = pt_update_op->prl,
};
- struct xe_pt *pt = vm->pt_root[tile->id];
+ struct xe_pt *pt = xe_vm_pt_root(vm, tile->id);
(void)xe_pt_walk_shared(&pt->base, pt->level, start, end,
&xe_walk.base);
@@ -2539,9 +2550,21 @@ int xe_pt_update_ops_prepare(struct xe_device *xe, struct xe_vma_ops *vops)
int id, err;
for_each_tile(tile, xe, id) {
+ struct xe_vm_pgtable_update_ops *pt_update_ops =
+ &vops->pt_update_ops[id];
+
if (!vops->pt_update_ops[id].num_ops)
continue;
+ if (pt_mirroring_disabled_for_tile(vops->vm, id)) {
+ struct xe_page_reclaim_list *prl = &pt_update_ops->prl;
+
+ /* Transfer root PT update ops PRL to current */
+ *prl = vops->pt_update_ops[0].prl;
+ xe_page_reclaim_entries_get(prl->entries);
+ continue;
+ }
+
err = __xe_pt_update_ops_prepare(tile, vops);
if (err)
return err;
@@ -2833,6 +2856,12 @@ xe_pt_update_ops_run(struct xe_device *xe, struct xe_vma_ops *vops)
struct xe_vm_pgtable_update_ops *pt_update_ops =
&vops->pt_update_ops[j];
+ if (pt_mirroring_disabled_for_tile(vm, j)) {
+ xe_tile_assert(tile, !get_current_op(pt_update_ops));
+ tile_mask |= BIT(tile->id);
+ continue;
+ }
+
for (i = 0; i < get_current_op(pt_update_ops); ++i) {
struct xe_vm_pgtable_update_op *pt_op =
to_pt_op(pt_update_ops, i);
@@ -2946,6 +2975,9 @@ void xe_pt_update_ops_abort(struct xe_device *xe, struct xe_vma_ops *vops)
&vops->pt_update_ops[id];
int i;
+ if (pt_mirroring_disabled_for_tile(vops->vm, id))
+ continue;
+
for (i = pt_update_ops->num_ops - 1; i >= 0; --i) {
struct xe_vm_pgtable_update_op *pt_op =
to_pt_op(pt_update_ops, i);
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index 0f542f47b9a8..f450e1c6f750 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -846,6 +846,14 @@ int xe_vm_rebind(struct xe_vm *vm, bool rebind_worker)
return err;
}
+static u8 adjust_rebind_tile_mask(struct xe_vm *vm, u8 tile_mask)
+{
+ if (vm->xe->info.has_pt_mirror)
+ return tile_mask;
+
+ return (0x1 << vm->xe->info.tile_count) - 1;
+}
+
struct dma_fence *xe_vma_rebind(struct xe_vm *vm, struct xe_vma *vma, u8 tile_mask)
{
struct dma_fence *fence = NULL;
@@ -862,7 +870,8 @@ struct dma_fence *xe_vma_rebind(struct xe_vm *vm, struct xe_vma *vma, u8 tile_ma
vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT |
XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
- err = xe_vm_ops_add_rebind(&vops, vma, tile_mask);
+ err = xe_vm_ops_add_rebind(&vops, vma,
+ adjust_rebind_tile_mask(vm, tile_mask));
if (err)
return ERR_PTR(err);
@@ -949,7 +958,8 @@ struct dma_fence *xe_vm_range_rebind(struct xe_vm *vm,
vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT |
XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
- err = xe_vm_ops_add_range_rebind(&vops, vma, range, tile_mask);
+ err = xe_vm_ops_add_range_rebind(&vops, vma, range,
+ adjust_rebind_tile_mask(vm, tile_mask));
if (err)
return ERR_PTR(err);
@@ -1739,7 +1749,8 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
for_each_tile(tile, xe, id) {
if (flags & XE_VM_FLAG_MIGRATION &&
- tile->id != XE_VM_FLAG_TILE_ID(flags))
+ tile->id != XE_VM_FLAG_TILE_ID(flags) &&
+ (vm->xe->info.has_pt_mirror || id))
continue;
vm->pt_root[id] = xe_pt_create(vm, tile, xe->info.vm_max_level,
@@ -2049,7 +2060,7 @@ struct xe_vm *xe_vm_lookup(struct xe_file *xef, u32 id)
u64 xe_vm_pdp4_descriptor(struct xe_vm *vm, struct xe_tile *tile)
{
- return vm->pt_ops->pde_encode_bo(vm->pt_root[tile->id]->bo, 0);
+ return vm->pt_ops->pde_encode_bo(xe_vm_pt_root(vm, tile->id)->bo, 0);
}
static struct xe_exec_queue *
@@ -5074,3 +5085,21 @@ void xe_vm_remove_exec_queue(struct xe_vm *vm, struct xe_exec_queue *q)
}
up_write(&vm->exec_queues.lock);
}
+
+/**
+ * xe_vm_pt_root() - Retrieve VM page-table root
+ * @vm: The VM.
+ * @tile_id: Tile ID
+ *
+ * Retrieve VM page-table root for a tile ID, used to abstract if PT mirroring is
+ * enabled across tiles.
+ *
+ * Return: VM page-table root for a tile ID
+ */
+struct xe_pt *xe_vm_pt_root(struct xe_vm *vm, u8 tile_id)
+{
+ if (vm->xe->info.has_pt_mirror)
+ return vm->pt_root[tile_id];
+
+ return vm->pt_root[0];
+}
diff --git a/drivers/gpu/drm/xe/xe_vm.h b/drivers/gpu/drm/xe/xe_vm.h
index c5b900f38ded..89c4f95984a2 100644
--- a/drivers/gpu/drm/xe/xe_vm.h
+++ b/drivers/gpu/drm/xe/xe_vm.h
@@ -436,4 +436,7 @@ static inline struct drm_exec *xe_vm_validation_exec(struct xe_vm *vm)
((READ_ONCE(tile_present) & ~READ_ONCE(tile_invalidated)) & BIT((tile)->id))
void xe_vma_mem_attr_copy(struct xe_vma_mem_attr *to, struct xe_vma_mem_attr *from);
+
+struct xe_pt *xe_vm_pt_root(struct xe_vm *vm, u8 tile_id);
+
#endif
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles
2026-09-25 4:53 ` [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
@ 2026-09-25 6:25 ` sashiko-bot
2026-09-25 7:21 ` Matthew Brost
2026-09-25 9:51 ` Francois Dugast
1 sibling, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 6:25 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] xe: wasted VRAM and forced cross-tile PT access for migration VMs when PT mirroring is disabled
--
commit ec37bcd236de3cdab70b39d0f31888be55640311
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: Add device flag to enable PT mirroring across tiles
This commit introduces a device flag to enable or disable page-table
mirroring across tiles. It defaults to enabled for PVC, aiming to allow
flexibility based on memory bandwidth needs.
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index 0f542f47b9a84..f450e1c6f7500 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
[ ... ]
> @@ -1739,7 +1749,8 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
>
> for_each_tile(tile, xe, id) {
> if (flags & XE_VM_FLAG_MIGRATION &&
> - tile->id != XE_VM_FLAG_TILE_ID(flags))
> + tile->id != XE_VM_FLAG_TILE_ID(flags) &&
> + (vm->xe->info.has_pt_mirror || id))
> continue;
>
> vm->pt_root[id] = xe_pt_create(vm, tile, xe->info.vm_max_level,
[Severity: Medium]
When has_pt_mirror is false, does this logic in xe_vm_create() allocate
page tables that will never be used?
For normal VMs (where flags & XE_VM_FLAG_MIGRATION is false), the condition
fails and we allocate pt_root[id] for all tiles. However, xe_vm_pt_root()
will unconditionally return pt_root[0] when has_pt_mirror is false, leaving
pt_root[1..N] unused and wasting VRAM.
Additionally, for a migration VM targeting a secondary tile (e.g., tile 1),
the condition tile->id != 1 && (has_pt_mirror || 0) evaluates to false when
id == 0, causing it to allocate both pt_root[0] and pt_root[1].
[ ... ]
> @@ -5074,3 +5085,21 @@ void xe_vm_remove_exec_queue(struct xe_vm *vm, struct xe_exec_queue *q)
> }
> up_write(&vm->exec_queues.lock);
> }
> +
> +/**
> + * xe_vm_pt_root() - Retrieve VM page-table root
> + * @vm: The VM.
> + * @tile_id: Tile ID
> + *
> + * Retrieve VM page-table root for a tile ID, used to abstract if PT mirroring is
> + * enabled across tiles.
> + *
> + * Return: VM page-table root for a tile ID
> + */
> +struct xe_pt *xe_vm_pt_root(struct xe_vm *vm, u8 tile_id)
> +{
> + if (vm->xe->info.has_pt_mirror)
> + return vm->pt_root[tile_id];
> +
> + return vm->pt_root[0];
> +}
[Severity: Medium]
Does returning vm->pt_root[0] here in xe_vm_pt_root() when has_pt_mirror
is false force cross-tile VRAM access for migration VMs on secondary tiles?
As noted above, a migration VM targeting a secondary tile will have its page
tables allocated in pt_root[1]. But since xe_vm_pt_root(vm, 1) will return
pt_root[0], the migration engine on tile 1 will be forced to walk page
tables physically located in tile 0's VRAM.
Can this cause a performance regression for migration operations on secondary
tiles?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=17
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles
2026-09-25 6:25 ` sashiko-bot
@ 2026-09-25 7:21 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 7:21 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 06:25:10AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] xe: wasted VRAM and forced cross-tile PT access for migration VMs when PT mirroring is disabled
> --
>
> commit ec37bcd236de3cdab70b39d0f31888be55640311
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Add device flag to enable PT mirroring across tiles
>
> This commit introduces a device flag to enable or disable page-table
> mirroring across tiles. It defaults to enabled for PVC, aiming to allow
> flexibility based on memory bandwidth needs.
>
> > diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> > index 0f542f47b9a84..f450e1c6f7500 100644
> > --- a/drivers/gpu/drm/xe/xe_vm.c
> > +++ b/drivers/gpu/drm/xe/xe_vm.c
> [ ... ]
> > @@ -1739,7 +1749,8 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
> >
> > for_each_tile(tile, xe, id) {
> > if (flags & XE_VM_FLAG_MIGRATION &&
> > - tile->id != XE_VM_FLAG_TILE_ID(flags))
> > + tile->id != XE_VM_FLAG_TILE_ID(flags) &&
> > + (vm->xe->info.has_pt_mirror || id))
> > continue;
> >
> > vm->pt_root[id] = xe_pt_create(vm, tile, xe->info.vm_max_level,
>
> [Severity: Medium]
> When has_pt_mirror is false, does this logic in xe_vm_create() allocate
> page tables that will never be used?
>
> For normal VMs (where flags & XE_VM_FLAG_MIGRATION is false), the condition
> fails and we allocate pt_root[id] for all tiles. However, xe_vm_pt_root()
> will unconditionally return pt_root[0] when has_pt_mirror is false, leaving
> pt_root[1..N] unused and wasting VRAM.
>
> Additionally, for a migration VM targeting a secondary tile (e.g., tile 1),
> the condition tile->id != 1 && (has_pt_mirror || 0) evaluates to false when
> id == 0, causing it to allocate both pt_root[0] and pt_root[1].
>
Yes, we over allocate 4k. I'm ok with this.
> [ ... ]
> > @@ -5074,3 +5085,21 @@ void xe_vm_remove_exec_queue(struct xe_vm *vm, struct xe_exec_queue *q)
> > }
> > up_write(&vm->exec_queues.lock);
> > }
> > +
> > +/**
> > + * xe_vm_pt_root() - Retrieve VM page-table root
> > + * @vm: The VM.
> > + * @tile_id: Tile ID
> > + *
> > + * Retrieve VM page-table root for a tile ID, used to abstract if PT mirroring is
> > + * enabled across tiles.
> > + *
> > + * Return: VM page-table root for a tile ID
> > + */
> > +struct xe_pt *xe_vm_pt_root(struct xe_vm *vm, u8 tile_id)
> > +{
> > + if (vm->xe->info.has_pt_mirror)
> > + return vm->pt_root[tile_id];
> > +
> > + return vm->pt_root[0];
> > +}
>
> [Severity: Medium]
> Does returning vm->pt_root[0] here in xe_vm_pt_root() when has_pt_mirror
> is false force cross-tile VRAM access for migration VMs on secondary tiles?
>
> As noted above, a migration VM targeting a secondary tile will have its page
> tables allocated in pt_root[1]. But since xe_vm_pt_root(vm, 1) will return
> pt_root[0], the migration engine on tile 1 will be forced to walk page
> tables physically located in tile 0's VRAM.
>
> Can this cause a performance regression for migration operations on secondary
> tiles?
>
!info.has_pt_mirror means VRAM access across tile is uniform perf or we'd mirror them.
Matt
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=17
^ permalink raw reply [flat|nested] 54+ messages in thread
* Re: [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles
2026-09-25 4:53 ` [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
2026-09-25 6:25 ` sashiko-bot
@ 2026-09-25 9:51 ` Francois Dugast
1 sibling, 0 replies; 54+ messages in thread
From: Francois Dugast @ 2026-09-25 9:51 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
On Thu, Sep 24, 2026 at 09:53:13PM -0700, Matthew Brost wrote:
> Some multi-tile devices may want to mirror page tables across tiles for
> memory-bandwidth reasons, while others may not. Add a device flag that
> allows enabling or disabling page-table mirroring across tiles.
>
> Setting the flag to true (the existing behavior) on PVC, but both modes
> have been tested and are working on PVC.
>
> Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Francois Dugast <francois.dugast@intel.com>
>
> ---
> v7:
> - Rework pt_mirroring_disabled_for_tile (Himal)
> ---
> drivers/gpu/drm/xe/xe_device_types.h | 2 ++
> drivers/gpu/drm/xe/xe_migrate.c | 5 ++--
> drivers/gpu/drm/xe/xe_pci.c | 2 ++
> drivers/gpu/drm/xe/xe_pci_types.h | 3 ++-
> drivers/gpu/drm/xe/xe_pt.c | 36 +++++++++++++++++++++++++--
> drivers/gpu/drm/xe/xe_vm.c | 37 +++++++++++++++++++++++++---
> drivers/gpu/drm/xe/xe_vm.h | 3 +++
> 7 files changed, 79 insertions(+), 9 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_device_types.h b/drivers/gpu/drm/xe/xe_device_types.h
> index aff2e5c8b656..cd9ed8ff2940 100644
> --- a/drivers/gpu/drm/xe/xe_device_types.h
> +++ b/drivers/gpu/drm/xe/xe_device_types.h
> @@ -229,6 +229,8 @@ struct xe_device {
> u8 has_usm:1;
> /** @info.has_64bit_timestamp: Device supports 64-bit timestamps */
> u8 has_64bit_timestamp:1;
> + /** @info.has_pt_mirror: Device has PT mirroring across tiles */
> + u8 has_pt_mirror:1;
> /** @info.is_dgfx: is discrete device */
> u8 is_dgfx:1;
> /** @info.needs_scratch: needs scratch page for oob prefetch to work */
> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> index f7e1a81434b2..471ae5741836 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
> @@ -246,7 +246,8 @@ static void xe_migrate_prepare_vm(struct xe_tile *tile, struct xe_migrate *m,
> struct xe_device *xe = tile_to_xe(tile);
> u16 pat_index = xe_cache_pat_idx(xe, XE_CACHE_WB);
> u8 id = tile->id;
> - u32 num_entries = NUM_PT_SLOTS, num_level = vm->pt_root[id]->level;
> + u32 num_entries = NUM_PT_SLOTS, num_level =
> + xe_vm_pt_root(vm, id)->level;
> #define VRAM_IDENTITY_MAP_PT_COUNT 4
> u32 num_setup = num_level + VRAM_IDENTITY_MAP_PT_COUNT;
> #undef VRAM_IDENTITY_MAP_PT_COUNT
> @@ -258,7 +259,7 @@ static void xe_migrate_prepare_vm(struct xe_tile *tile, struct xe_migrate *m,
> u64 l1_pt_ofs = xe_bo_size(bo) - 5 * XE_PAGE_SIZE;
>
> entry = vm->pt_ops->pde_encode_bo(bo, l1_pt_ofs);
> - xe_pt_write(xe, &vm->pt_root[id]->bo->vmap, 0, entry);
> + xe_pt_write(xe, &xe_vm_pt_root(vm, id)->bo->vmap, 0, entry);
>
> map_ofs = (num_entries - num_setup) * XE_PAGE_SIZE;
>
> diff --git a/drivers/gpu/drm/xe/xe_pci.c b/drivers/gpu/drm/xe/xe_pci.c
> index d64e873d9e71..28f37e034e02 100644
> --- a/drivers/gpu/drm/xe/xe_pci.c
> +++ b/drivers/gpu/drm/xe/xe_pci.c
> @@ -371,6 +371,7 @@ static const __maybe_unused struct xe_device_desc pvc_desc = {
> .has_display = false,
> .has_drm_ras = true,
> .has_gsc_nvm = 1,
> + .has_pt_mirror = 1,
> .has_heci_gscfi = 1,
> .max_gt_per_tile = 1,
> .max_remote_tiles = 1,
> @@ -807,6 +808,7 @@ static int xe_info_init_early(struct xe_device *xe,
> xe->info.has_mert = desc->has_mert;
> xe->info.has_page_reclaim_hw_assist = desc->has_page_reclaim_hw_assist;
> xe->info.has_pre_prod_wa = desc->has_pre_prod_wa;
> + xe->info.has_pt_mirror = desc->has_pt_mirror;
> xe->info.has_pxp = desc->has_pxp;
> xe->info.has_soc_remapper_sysctrl = desc->has_soc_remapper_sysctrl;
> xe->info.has_soc_remapper_telem = desc->has_soc_remapper_telem;
> diff --git a/drivers/gpu/drm/xe/xe_pci_types.h b/drivers/gpu/drm/xe/xe_pci_types.h
> index fed509ff601e..71068cdb3558 100644
> --- a/drivers/gpu/drm/xe/xe_pci_types.h
> +++ b/drivers/gpu/drm/xe/xe_pci_types.h
> @@ -52,8 +52,9 @@ struct xe_device_desc {
> u8 has_mbx_power_limits:1;
> u8 has_mbx_thermal_info:1;
> u8 has_mert:1;
> - u8 has_pre_prod_wa:1;
> u8 has_page_reclaim_hw_assist:1;
> + u8 has_pre_prod_wa:1;
> + u8 has_pt_mirror:1;
> u8 has_pxp:1;
> u8 has_soc_remapper_sysctrl:1;
> u8 has_soc_remapper_telem:1;
> diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
> index fbefcfc52dca..dcef29b56d63 100644
> --- a/drivers/gpu/drm/xe/xe_pt.c
> +++ b/drivers/gpu/drm/xe/xe_pt.c
> @@ -804,7 +804,7 @@ xe_pt_stage_bind(struct xe_tile *tile, struct xe_vma *vma,
> .wupd.entries = entries,
> .clear_pt = clear_pt,
> };
> - struct xe_pt *pt = vm->pt_root[tile->id];
> + struct xe_pt *pt = xe_vm_pt_root(vm, tile->id);
> int ret;
> bool is_purged = false;
>
> @@ -1005,6 +1005,11 @@ static int xe_pt_zap_ptes_entry(struct xe_ptw *parent, pgoff_t offset,
> return 0;
> }
>
> +static bool pt_mirroring_disabled_for_tile(struct xe_vm *vm, u8 tile_id)
> +{
> + return !vm->xe->info.has_pt_mirror && tile_id;
> +}
> +
> static const struct xe_pt_walk_ops xe_pt_zap_ptes_ops = {
> .pt_entry = xe_pt_zap_ptes_entry,
> };
> @@ -1046,6 +1051,9 @@ bool xe_pt_zap_ptes(struct xe_tile *tile, struct xe_vma *vma)
> if (!(pt_mask & BIT(tile->id)))
> return false;
>
> + if (pt_mirroring_disabled_for_tile(xe_vma_vm(vma), tile->id))
> + return true;
> +
> (void)xe_pt_walk_shared(&pt->base, pt->level, xe_vma_start(vma),
> xe_vma_end(vma), &xe_walk.base);
>
> @@ -1098,6 +1106,9 @@ bool xe_pt_zap_ptes_range(struct xe_tile *tile, struct xe_vm *vm,
> if (!(pt_mask & BIT(tile->id)))
> return false;
>
> + if (pt_mirroring_disabled_for_tile(vm, tile->id))
> + return true;
> +
> (void)xe_pt_walk_shared(&pt->base, pt->level, xe_svm_range_start(range),
> xe_svm_range_end(range), &xe_walk.base);
>
> @@ -1987,7 +1998,7 @@ static unsigned int xe_pt_stage_unbind(struct xe_tile *tile,
> .wupd.entries = entries,
> .prl = pt_update_op->prl,
> };
> - struct xe_pt *pt = vm->pt_root[tile->id];
> + struct xe_pt *pt = xe_vm_pt_root(vm, tile->id);
>
> (void)xe_pt_walk_shared(&pt->base, pt->level, start, end,
> &xe_walk.base);
> @@ -2539,9 +2550,21 @@ int xe_pt_update_ops_prepare(struct xe_device *xe, struct xe_vma_ops *vops)
> int id, err;
>
> for_each_tile(tile, xe, id) {
> + struct xe_vm_pgtable_update_ops *pt_update_ops =
> + &vops->pt_update_ops[id];
> +
> if (!vops->pt_update_ops[id].num_ops)
> continue;
>
> + if (pt_mirroring_disabled_for_tile(vops->vm, id)) {
> + struct xe_page_reclaim_list *prl = &pt_update_ops->prl;
> +
> + /* Transfer root PT update ops PRL to current */
> + *prl = vops->pt_update_ops[0].prl;
> + xe_page_reclaim_entries_get(prl->entries);
> + continue;
> + }
> +
> err = __xe_pt_update_ops_prepare(tile, vops);
> if (err)
> return err;
> @@ -2833,6 +2856,12 @@ xe_pt_update_ops_run(struct xe_device *xe, struct xe_vma_ops *vops)
> struct xe_vm_pgtable_update_ops *pt_update_ops =
> &vops->pt_update_ops[j];
>
> + if (pt_mirroring_disabled_for_tile(vm, j)) {
> + xe_tile_assert(tile, !get_current_op(pt_update_ops));
> + tile_mask |= BIT(tile->id);
> + continue;
> + }
> +
> for (i = 0; i < get_current_op(pt_update_ops); ++i) {
> struct xe_vm_pgtable_update_op *pt_op =
> to_pt_op(pt_update_ops, i);
> @@ -2946,6 +2975,9 @@ void xe_pt_update_ops_abort(struct xe_device *xe, struct xe_vma_ops *vops)
> &vops->pt_update_ops[id];
> int i;
>
> + if (pt_mirroring_disabled_for_tile(vops->vm, id))
> + continue;
> +
> for (i = pt_update_ops->num_ops - 1; i >= 0; --i) {
> struct xe_vm_pgtable_update_op *pt_op =
> to_pt_op(pt_update_ops, i);
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index 0f542f47b9a8..f450e1c6f750 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
> @@ -846,6 +846,14 @@ int xe_vm_rebind(struct xe_vm *vm, bool rebind_worker)
> return err;
> }
>
> +static u8 adjust_rebind_tile_mask(struct xe_vm *vm, u8 tile_mask)
> +{
> + if (vm->xe->info.has_pt_mirror)
> + return tile_mask;
> +
> + return (0x1 << vm->xe->info.tile_count) - 1;
> +}
> +
> struct dma_fence *xe_vma_rebind(struct xe_vm *vm, struct xe_vma *vma, u8 tile_mask)
> {
> struct dma_fence *fence = NULL;
> @@ -862,7 +870,8 @@ struct dma_fence *xe_vma_rebind(struct xe_vm *vm, struct xe_vma *vma, u8 tile_ma
> vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT |
> XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
>
> - err = xe_vm_ops_add_rebind(&vops, vma, tile_mask);
> + err = xe_vm_ops_add_rebind(&vops, vma,
> + adjust_rebind_tile_mask(vm, tile_mask));
> if (err)
> return ERR_PTR(err);
>
> @@ -949,7 +958,8 @@ struct dma_fence *xe_vm_range_rebind(struct xe_vm *vm,
> vops.flags |= XE_VMA_OPS_FLAG_SKIP_TLB_WAIT |
> XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP;
>
> - err = xe_vm_ops_add_range_rebind(&vops, vma, range, tile_mask);
> + err = xe_vm_ops_add_range_rebind(&vops, vma, range,
> + adjust_rebind_tile_mask(vm, tile_mask));
> if (err)
> return ERR_PTR(err);
>
> @@ -1739,7 +1749,8 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
>
> for_each_tile(tile, xe, id) {
> if (flags & XE_VM_FLAG_MIGRATION &&
> - tile->id != XE_VM_FLAG_TILE_ID(flags))
> + tile->id != XE_VM_FLAG_TILE_ID(flags) &&
> + (vm->xe->info.has_pt_mirror || id))
> continue;
>
> vm->pt_root[id] = xe_pt_create(vm, tile, xe->info.vm_max_level,
> @@ -2049,7 +2060,7 @@ struct xe_vm *xe_vm_lookup(struct xe_file *xef, u32 id)
>
> u64 xe_vm_pdp4_descriptor(struct xe_vm *vm, struct xe_tile *tile)
> {
> - return vm->pt_ops->pde_encode_bo(vm->pt_root[tile->id]->bo, 0);
> + return vm->pt_ops->pde_encode_bo(xe_vm_pt_root(vm, tile->id)->bo, 0);
> }
>
> static struct xe_exec_queue *
> @@ -5074,3 +5085,21 @@ void xe_vm_remove_exec_queue(struct xe_vm *vm, struct xe_exec_queue *q)
> }
> up_write(&vm->exec_queues.lock);
> }
> +
> +/**
> + * xe_vm_pt_root() - Retrieve VM page-table root
> + * @vm: The VM.
> + * @tile_id: Tile ID
> + *
> + * Retrieve VM page-table root for a tile ID, used to abstract if PT mirroring is
> + * enabled across tiles.
> + *
> + * Return: VM page-table root for a tile ID
> + */
> +struct xe_pt *xe_vm_pt_root(struct xe_vm *vm, u8 tile_id)
> +{
> + if (vm->xe->info.has_pt_mirror)
> + return vm->pt_root[tile_id];
> +
> + return vm->pt_root[0];
> +}
> diff --git a/drivers/gpu/drm/xe/xe_vm.h b/drivers/gpu/drm/xe/xe_vm.h
> index c5b900f38ded..89c4f95984a2 100644
> --- a/drivers/gpu/drm/xe/xe_vm.h
> +++ b/drivers/gpu/drm/xe/xe_vm.h
> @@ -436,4 +436,7 @@ static inline struct drm_exec *xe_vm_validation_exec(struct xe_vm *vm)
> ((READ_ONCE(tile_present) & ~READ_ONCE(tile_invalidated)) & BIT((tile)->id))
>
> void xe_vma_mem_attr_copy(struct xe_vma_mem_attr *to, struct xe_vma_mem_attr *from);
> +
> +struct xe_pt *xe_vm_pt_root(struct xe_vm *vm, u8 tile_id);
> +
> #endif
> --
> 2.34.1
>
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 18/24] drm/xe: Add ULLS support to LRC
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (16 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
` (10 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Himal Prasad Ghimiray
Define memory layout for ULLS semaphores stored in LRC memory. Add
support functions to return GGTT address and set semaphore based on a
job's seqno.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
drivers/gpu/drm/xe/xe_lrc.c | 73 +++++++++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_lrc.h | 4 ++
drivers/gpu/drm/xe/xe_lrc_types.h | 4 ++
3 files changed, 81 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_lrc.c b/drivers/gpu/drm/xe/xe_lrc.c
index f1cf1463f1b2..bf5c2351aed3 100644
--- a/drivers/gpu/drm/xe/xe_lrc.c
+++ b/drivers/gpu/drm/xe/xe_lrc.c
@@ -706,6 +706,7 @@ u32 xe_lrc_pphwsp_offset(struct xe_lrc *lrc)
#define LRC_CTX_JOB_TIMESTAMP_OFFSET 512
#define LRC_ENGINE_ID_PPHWSP_OFFSET 1024
#define LRC_PARALLEL_PPHWSP_OFFSET 2048
+#define LRC_ULLS_PPHWSP_OFFSET 2048 /* Mutually exclusive with parallel */
#define LRC_SEQNO_OFFSET 0
#define LRC_START_SEQNO_OFFSET (LRC_SEQNO_OFFSET + 8)
@@ -768,6 +769,12 @@ static inline u32 __xe_lrc_engine_id_offset(struct xe_lrc *lrc)
return xe_lrc_pphwsp_offset(lrc) + LRC_ENGINE_ID_PPHWSP_OFFSET;
}
+static u32 __xe_lrc_ulls_offset(struct xe_lrc *lrc)
+{
+ /* The ulls is stored in the driver-defined portion of PPHWSP */
+ return xe_lrc_pphwsp_offset(lrc) + LRC_ULLS_PPHWSP_OFFSET;
+}
+
static u32 __xe_lrc_ctx_timestamp_offset(struct xe_lrc *lrc)
{
return __xe_lrc_regs_offset(lrc) + CTX_TIMESTAMP * sizeof(u32);
@@ -835,6 +842,7 @@ DECL_MAP_ADDR_HELPERS(ctx_job_timestamp, lrc->bo)
DECL_MAP_ADDR_HELPERS(ctx_timestamp, lrc->bo)
DECL_MAP_ADDR_HELPERS(ctx_timestamp_udw, lrc->bo)
DECL_MAP_ADDR_HELPERS(parallel, lrc->bo)
+DECL_MAP_ADDR_HELPERS(ulls, lrc->bo)
DECL_MAP_ADDR_HELPERS(indirect_ring, lrc->bo)
DECL_MAP_ADDR_HELPERS(engine_id, lrc->bo)
DECL_MAP_ADDR_HELPERS(queue_timestamp, lrc->bo)
@@ -1805,6 +1813,26 @@ void xe_lrc_set_ring_tail(struct xe_lrc *lrc, u32 tail)
xe_lrc_write_ctx_reg(lrc, CTX_RING_TAIL, tail);
}
+/**
+ * xe_lrc_ring_tail_ggtt_addr() - Saved ring tail GGTT address
+ * @lrc: Pointer to the lrc.
+ *
+ * GGTT address of the ring tail as saved for this context - in the indirect
+ * ring state on platforms which have it, otherwise in the context image.
+ * This is what a context restore loads the tail register from, so a ring
+ * which advances the tail register itself must keep this in sync.
+ *
+ * Returns: saved ring tail GGTT address
+ */
+u32 xe_lrc_ring_tail_ggtt_addr(struct xe_lrc *lrc)
+{
+ if (xe_lrc_has_indirect_ring_state(lrc))
+ return __xe_lrc_indirect_ring_ggtt_addr(lrc) +
+ INDIRECT_CTX_RING_TAIL * sizeof(u32);
+
+ return __xe_lrc_regs_ggtt_addr(lrc) + CTX_RING_TAIL * sizeof(u32);
+}
+
u32 xe_lrc_ring_tail(struct xe_lrc *lrc)
{
if (xe_lrc_has_indirect_ring_state(lrc))
@@ -1984,6 +2012,51 @@ static u32 xe_lrc_engine_id(struct xe_lrc *lrc)
return xe_map_read32(xe, &map);
}
+#define semaphore_offset(seqno) \
+ (sizeof(u32) * ((seqno) % LRC_MIGRATION_ULLS_SEMAPHORE_COUNT))
+
+/**
+ * xe_lrc_ulls_semaphore_ggtt_addr() - ULLS semaphore GGTT address
+ * @lrc: Pointer to the lrc.
+ * @seqno: seqno of current job.
+ *
+ * Calculate ULLS semaphore GGTT address based on input seqno
+ *
+ * Returns: ULLS semaphore GGTT address
+ */
+u32 xe_lrc_ulls_semaphore_ggtt_addr(struct xe_lrc *lrc, u32 seqno)
+{
+ xe_assert(lrc_to_xe(lrc), semaphore_offset(seqno) <
+ LRC_PPHWSP_SIZE - LRC_ULLS_PPHWSP_OFFSET);
+
+ return __xe_lrc_ulls_ggtt_addr(lrc) + semaphore_offset(seqno);
+}
+
+/**
+ * xe_lrc_set_ulls_semaphore() - Set ULLS semaphore
+ * @lrc: Pointer to the lrc.
+ * @seqno: seqno of current job.
+ *
+ * Set ULLS semaphore based on input seqno
+ */
+void xe_lrc_set_ulls_semaphore(struct xe_lrc *lrc, u32 seqno)
+{
+ struct xe_device *xe = lrc_to_xe(lrc);
+ struct iosys_map map = __xe_lrc_ulls_map(lrc);
+
+ xe_assert(xe, semaphore_offset(seqno) <
+ LRC_PPHWSP_SIZE - LRC_ULLS_PPHWSP_OFFSET);
+
+ /*
+ * The ring contents this semaphore releases are ordered by the
+ * xe_device_wmb() at the end of xe_lrc_write_ring().
+ */
+ iosys_map_incr(&map, semaphore_offset(seqno));
+ xe_map_write32(xe, &map, LRC_MIGRATION_ULLS_SEMAPHORE_SIGNAL);
+
+ xe_device_wmb(xe); /* Flush write to hardware */
+}
+
static int instr_dw(u32 cmd_header)
{
/* GFXPIPE "SINGLE_DW" opcodes are a single dword */
diff --git a/drivers/gpu/drm/xe/xe_lrc.h b/drivers/gpu/drm/xe/xe_lrc.h
index a8ff4e59a1f4..f2e829e62168 100644
--- a/drivers/gpu/drm/xe/xe_lrc.h
+++ b/drivers/gpu/drm/xe/xe_lrc.h
@@ -112,6 +112,7 @@ u32 xe_lrc_regs_offset(struct xe_lrc *lrc);
void xe_lrc_set_ring_tail(struct xe_lrc *lrc, u32 tail);
u32 xe_lrc_ring_tail(struct xe_lrc *lrc);
+u32 xe_lrc_ring_tail_ggtt_addr(struct xe_lrc *lrc);
void xe_lrc_set_ring_head(struct xe_lrc *lrc, u32 head);
u32 xe_lrc_ring_head(struct xe_lrc *lrc);
u32 xe_lrc_ring_space(struct xe_lrc *lrc);
@@ -127,6 +128,9 @@ void xe_default_lrc_update_memirq_regs_with_address(struct xe_hw_engine *hwe);
void xe_lrc_update_memirq_regs_with_address(struct xe_lrc *lrc, struct xe_hw_engine *hwe,
u32 *regs);
+u32 xe_lrc_ulls_semaphore_ggtt_addr(struct xe_lrc *lrc, u32 seqno);
+void xe_lrc_set_ulls_semaphore(struct xe_lrc *lrc, u32 seqno);
+
u32 xe_lrc_read_ctx_reg(struct xe_lrc *lrc, int reg_nr);
void xe_lrc_write_ctx_reg(struct xe_lrc *lrc, int reg_nr, u32 val);
diff --git a/drivers/gpu/drm/xe/xe_lrc_types.h b/drivers/gpu/drm/xe/xe_lrc_types.h
index 53ef48feebfc..b18daf1092e6 100644
--- a/drivers/gpu/drm/xe/xe_lrc_types.h
+++ b/drivers/gpu/drm/xe/xe_lrc_types.h
@@ -12,6 +12,10 @@
struct xe_bo;
+#define LRC_MIGRATION_ULLS_SEMAPHORE_COUNT 64 /* Must be pow2 */
+#define LRC_MIGRATION_ULLS_SEMAPHORE_CLEAR 0
+#define LRC_MIGRATION_ULLS_SEMAPHORE_SIGNAL 1
+
/**
* struct xe_lrc - Logical ring context (LRC) and submission ring object
*/
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (17 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 18/24] drm/xe: Add ULLS support to LRC Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 6:34 ` sashiko-bot
2026-09-25 4:53 ` [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
` (9 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe
Add function to enter ULLS mode for migration job and delayed worker to
exit (power saving). ULLS mode expected to entered upon page fault or
SVM prefetch. ULLS mode exit delay is currently set to 5ms.
ULLS mode only support on DGFX and USM platforms where a hardware engine
is reserved for migrations jobs. When in ULLS mode, set several flags on
migration jobs so submission backend / ring ops can properly submit in
ULLS mode.
Upon ULLS mode enter, send a job trigger waiting a semphore pipling
initial GuC / HW conetxt switch.
Upon ULLS mode exit, send a job to trigger that current ULLS
semaphore so the ring can be taken off the hardware.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
---
v7:
- Style cleanups
- xe_migrate_ulls_exit don't wiat for dma-fence (Sashiko)
---
drivers/gpu/drm/xe/xe_exec_queue.c | 5 +-
drivers/gpu/drm/xe/xe_exec_queue.h | 2 +-
drivers/gpu/drm/xe/xe_migrate.c | 198 ++++++++++++++++++++++--
drivers/gpu/drm/xe/xe_migrate.h | 2 +
drivers/gpu/drm/xe/xe_pt.c | 2 +-
drivers/gpu/drm/xe/xe_sched_job.h | 56 +++++++
drivers/gpu/drm/xe/xe_sched_job_types.h | 19 +++
drivers/gpu/drm/xe/xe_vm.c | 2 +-
8 files changed, 265 insertions(+), 21 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.c b/drivers/gpu/drm/xe/xe_exec_queue.c
index f18229f65ad2..9a8efa2bf2bf 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.c
+++ b/drivers/gpu/drm/xe/xe_exec_queue.c
@@ -1546,6 +1546,7 @@ bool xe_exec_queue_is_lr(struct xe_exec_queue *q)
/**
* xe_exec_queue_is_idle() - Whether an exec_queue is idle.
* @q: The exec_queue
+ * @extra_jobs: Extra jobs on the queue
*
* FIXME: Need to determine what to use as the short-lived
* timeline lock for the exec_queues, so that the return value
@@ -1557,9 +1558,9 @@ bool xe_exec_queue_is_lr(struct xe_exec_queue *q)
*
* Return: True if the exec_queue is idle, false otherwise.
*/
-bool xe_exec_queue_is_idle(struct xe_exec_queue *q)
+bool xe_exec_queue_is_idle(struct xe_exec_queue *q, int extra_jobs)
{
- return !atomic_read(&q->job_cnt);
+ return !(atomic_read(&q->job_cnt) - extra_jobs);
}
/**
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.h b/drivers/gpu/drm/xe/xe_exec_queue.h
index b02a390ba989..e8963f85cabd 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.h
+++ b/drivers/gpu/drm/xe/xe_exec_queue.h
@@ -116,7 +116,7 @@ static inline struct xe_exec_queue *xe_exec_queue_multi_queue_primary(struct xe_
bool xe_exec_queue_is_lr(struct xe_exec_queue *q);
-bool xe_exec_queue_is_idle(struct xe_exec_queue *q);
+bool xe_exec_queue_is_idle(struct xe_exec_queue *q, int extra_jobs);
void xe_exec_queue_kill(struct xe_exec_queue *q);
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 471ae5741836..d7d13a25cdb9 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -8,6 +8,7 @@
#include <linux/bitfield.h>
#include <linux/sizes.h>
+#include <drm/drm_drv.h>
#include <drm/drm_managed.h>
#include <drm/drm_pagemap.h>
#include <drm/ttm/ttm_tt.h>
@@ -32,6 +33,7 @@
#include "xe_mem_pool.h"
#include "xe_mocs.h"
#include "xe_pat.h"
+#include "xe_pm.h"
#include "xe_printk.h"
#include "xe_pt.h"
#include "xe_res_cursor.h"
@@ -77,6 +79,14 @@ struct xe_migrate {
struct dma_fence *fence;
/** @min_chunk_size: For dgfx, Minimum chunk size */
u64 min_chunk_size;
+ /** @ulls: ULLS support */
+ struct {
+ /** @ulls.enabled: ULLS is enabled, protected by job_mutex */
+ bool enabled;
+#define ULLS_EXIT_JIFFIES msecs_to_jiffies(5)
+ /** @ulls.exit_work: ULLS exit worker */
+ struct delayed_work exit_work;
+ } ulls;
};
#define MAX_PREEMPTDISABLE_TRANSFER SZ_8M /* Around 1ms. */
@@ -95,9 +105,30 @@ struct xe_migrate {
*/
#define MAX_PTE_PER_SDI 0x1FEU
+static bool xe_migrate_ulls_enabled(struct xe_migrate *m)
+{
+ lockdep_assert_held(&m->job_mutex);
+ return m->ulls.enabled;
+}
+
+static void xe_migrate_ulls_toggle_enable(struct xe_migrate *m, bool enabled)
+{
+ lockdep_assert_held(&m->job_mutex);
+ m->ulls.enabled = enabled;
+}
+
static void xe_migrate_fini(void *arg)
{
struct xe_migrate *m = arg;
+ struct xe_device *xe = tile_to_xe(m->tile);
+
+ disable_delayed_work_sync(&m->ulls.exit_work);
+ scoped_guard(mutex, &m->job_mutex) {
+ if (xe_migrate_ulls_enabled(m)) {
+ xe_pm_runtime_put(xe);
+ xe_migrate_ulls_toggle_enable(m, false);
+ }
+ }
xe_vm_lock(m->q->vm, false);
xe_bo_unpin(m->pt_bo);
@@ -448,6 +479,150 @@ static int xe_migrate_lock_prepare_vm(struct xe_tile *tile, struct xe_migrate *m
return err;
}
+static struct dma_fence *__xe_migrate_job_push(struct xe_migrate *m,
+ struct xe_sched_job *job,
+ enum xe_ulls_state ulls)
+{
+ struct dma_fence *fence;
+
+ lockdep_assert_held(&m->job_mutex);
+ xe_tile_assert(m->tile, m->q == job->q);
+
+ job->ulls = ulls;
+ xe_sched_job_arm(job);
+ fence = dma_fence_get(&job->drm.s_fence->finished);
+ xe_sched_job_push(job);
+
+ return fence;
+}
+
+/*
+ * Arm and push a migration job, tagging it as a ULLS job and deferring the
+ * ULLS exit while ULLS mode is active.
+ *
+ * Returns a reference to the job's finished fence.
+ */
+static struct dma_fence *xe_migrate_job_push(struct xe_migrate *m,
+ struct xe_sched_job *job)
+{
+ enum xe_ulls_state ulls = ULLS_NONE;
+
+ lockdep_assert_held(&m->job_mutex);
+
+ if (xe_migrate_ulls_enabled(m)) {
+ ulls = ULLS_ACTIVE;
+ mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
+ ULLS_EXIT_JIFFIES);
+ }
+
+ return __xe_migrate_job_push(m, job, ulls);
+}
+
+/**
+ * xe_migrate_ulls_enter() - Enter ULLS mode
+ * @m: The migration context.
+ *
+ * If DGFX, enter ULLS mode bypassing GuC / HW context switches by utilizing
+ * semaphore and continuously running batches.
+ */
+void xe_migrate_ulls_enter(struct xe_migrate *m)
+{
+ struct xe_device *xe = tile_to_xe(m->tile);
+ struct xe_sched_job *job = NULL;
+ u64 batch_addr[2] = { 0, 0 };
+ bool alloc = false;
+
+ xe_assert(xe, xe->info.has_usm);
+
+ if (!IS_DGFX(xe))
+ return;
+
+job_alloc:
+ if (alloc) {
+ /*
+ * Must be done outside job_mutex as that lock is tainted with
+ * reclaim.
+ */
+ job = xe_sched_job_create(m->q, batch_addr);
+ if (WARN_ON_ONCE(IS_ERR(job)))
+ return; /* Not fatal */
+ }
+
+ mutex_lock(&m->job_mutex);
+ if (!xe_migrate_ulls_enabled(m)) {
+ struct dma_fence *fence;
+
+ if (!job) {
+ alloc = true;
+ mutex_unlock(&m->job_mutex);
+ goto job_alloc;
+ }
+
+ /* Pairs with PM put on ULLS exit */
+ xe_pm_runtime_get_noresume(xe);
+
+ xe_sched_job_get(job);
+ fence = __xe_migrate_job_push(m, job, ULLS_ENTER);
+ dma_fence_put(fence);
+
+ xe_dbg(xe, "Migrate ULLS mode enter");
+ xe_migrate_ulls_toggle_enable(m, true);
+ }
+ if (job)
+ xe_sched_job_put(job);
+ if (xe_migrate_ulls_enabled(m))
+ mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
+ ULLS_EXIT_JIFFIES);
+ mutex_unlock(&m->job_mutex);
+}
+
+static void xe_migrate_ulls_exit(struct work_struct *work)
+{
+ struct xe_migrate *m = container_of(work, struct xe_migrate,
+ ulls.exit_work.work);
+ struct xe_device *xe = tile_to_xe(m->tile);
+ struct xe_sched_job *job = NULL;
+ struct dma_fence *fence = NULL;
+ u64 batch_addr[2] = { 0, 0 };
+ int idx;
+
+ xe_assert(xe, m->ulls.enabled);
+
+ if (!drm_dev_enter(&xe->drm, &idx))
+ return;
+
+ /*
+ * Must be done outside job_mutex as that lock is tainted with
+ * reclaim and must be done holding a pm ref.
+ */
+ job = xe_sched_job_create(m->q, batch_addr);
+ if (WARN_ON_ONCE(IS_ERR(job))) {
+ drm_dev_exit(idx);
+ mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
+ ULLS_EXIT_JIFFIES);
+ return; /* Not fatal */
+ }
+
+ scoped_guard(mutex, &m->job_mutex) {
+ if (xe_exec_queue_is_idle(m->q, 1)) {
+ fence = __xe_migrate_job_push(m, job, ULLS_EXIT);
+ dma_fence_put(fence);
+
+ xe_pm_runtime_put(xe); /* Pairs with PM get in enter */
+ xe_migrate_ulls_toggle_enable(m, false);
+ cancel_delayed_work(&m->ulls.exit_work);
+
+ xe_dbg(xe, "Migrate ULLS mode exit");
+ } else {
+ xe_sched_job_put(job);
+ mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
+ ULLS_EXIT_JIFFIES);
+ }
+ }
+
+ drm_dev_exit(idx);
+}
+
/**
* xe_migrate_init() - Initialize a migrate context
* @m: The migration context
@@ -506,6 +681,8 @@ int xe_migrate_init(struct xe_migrate *m)
might_lock(&m->job_mutex);
fs_reclaim_release(GFP_KERNEL);
+ INIT_DELAYED_WORK(&m->ulls.exit_work, xe_migrate_ulls_exit);
+
err = devm_add_action_or_reset(xe->drm.dev, xe_migrate_fini, m);
if (err)
return err;
@@ -1033,10 +1210,8 @@ static struct dma_fence *__xe_migrate_copy(struct xe_migrate *m,
}
mutex_lock(&m->job_mutex);
- xe_sched_job_arm(job);
dma_fence_put(fence);
- fence = dma_fence_get(&job->drm.s_fence->finished);
- xe_sched_job_push(job);
+ fence = xe_migrate_job_push(m, job);
dma_fence_put(m->fence);
m->fence = dma_fence_get(fence);
@@ -1464,10 +1639,8 @@ struct dma_fence *xe_migrate_vram_copy_chunk(struct xe_bo *vram_bo, u64 vram_off
DMA_RESV_USAGE_BOOKKEEP));
scoped_guard(mutex, &m->job_mutex) {
- xe_sched_job_arm(job);
dma_fence_put(fence);
- fence = dma_fence_get(&job->drm.s_fence->finished);
- xe_sched_job_push(job);
+ fence = xe_migrate_job_push(m, job);
dma_fence_put(m->fence);
m->fence = dma_fence_get(fence);
@@ -1701,10 +1874,8 @@ struct dma_fence *xe_migrate_clear(struct xe_migrate *m,
}
mutex_lock(&m->job_mutex);
- xe_sched_job_arm(job);
dma_fence_put(fence);
- fence = dma_fence_get(&job->drm.s_fence->finished);
- xe_sched_job_push(job);
+ fence = xe_migrate_job_push(m, job);
dma_fence_put(m->fence);
m->fence = dma_fence_get(fence);
@@ -1980,9 +2151,7 @@ static struct dma_fence *xe_migrate_vram(struct xe_migrate *m,
}
mutex_lock(&m->job_mutex);
- xe_sched_job_arm(job);
- fence = dma_fence_get(&job->drm.s_fence->finished);
- xe_sched_job_push(job);
+ fence = xe_migrate_job_push(m, job);
dma_fence_put(m->fence);
m->fence = dma_fence_get(fence);
@@ -2299,10 +2468,7 @@ int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
xe_sched_job_add_migrate_flush(job, MI_FLUSH_DW_CCS);
mutex_lock(&m->job_mutex);
- xe_sched_job_arm(job);
-
- fence = dma_fence_get(&job->drm.s_fence->finished);
- xe_sched_job_push(job);
+ fence = xe_migrate_job_push(m, job);
mutex_unlock(&m->job_mutex);
dma_fence_wait(fence, false);
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index fa381ec36ef1..2e8be10fdb71 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -97,4 +97,6 @@ int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
bool write_to_ccs);
#endif
+void xe_migrate_ulls_enter(struct xe_migrate *m);
+
#endif
diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
index dcef29b56d63..6ad4bbaa1860 100644
--- a/drivers/gpu/drm/xe/xe_pt.c
+++ b/drivers/gpu/drm/xe/xe_pt.c
@@ -1426,7 +1426,7 @@ static int xe_pt_vm_dependencies(struct xe_sched_job *job,
if (!job && !no_in_syncs(vops->syncs, vops->num_syncs))
return -ETIME;
- if (!job && !xe_exec_queue_is_idle(vops->q))
+ if (!job && !xe_exec_queue_is_idle(vops->q, 0))
return -ETIME;
if (vops->flags & (XE_VMA_OPS_FLAG_WAIT_VM_BOOKKEEP |
diff --git a/drivers/gpu/drm/xe/xe_sched_job.h b/drivers/gpu/drm/xe/xe_sched_job.h
index 1c1cb44216c3..bbf5e3864ba3 100644
--- a/drivers/gpu/drm/xe/xe_sched_job.h
+++ b/drivers/gpu/drm/xe/xe_sched_job.h
@@ -83,6 +83,62 @@ xe_sched_job_add_migrate_flush(struct xe_sched_job *job, u32 flags)
job->migrate_flush_flags = flags;
}
+/**
+ * xe_sched_job_is_ulls() - Is a ULLS job
+ * @job: Xe schedule job object
+ *
+ * Return: True if @job is submitted as part of a ULLS sequence, False
+ * otherwise.
+ */
+static inline bool xe_sched_job_is_ulls(struct xe_sched_job *job)
+{
+ return job->ulls != ULLS_NONE;
+}
+
+/**
+ * xe_sched_job_ulls_has_batch() - Does a job carry batch buffers
+ * @job: Xe schedule job object
+ *
+ * The ULLS jobs which enter and exit ULLS mode exist only to move the
+ * migration context on and off the hardware, and carry no batch buffers.
+ *
+ * Return: True if @job carries batch buffers, False otherwise.
+ */
+static inline bool xe_sched_job_ulls_has_batch(struct xe_sched_job *job)
+{
+ return job->ulls == ULLS_NONE || job->ulls == ULLS_ACTIVE;
+}
+
+/**
+ * xe_sched_job_ulls_parks() - Does a job park the engine for its successor
+ * @job: Xe schedule job object
+ *
+ * A ULLS job which is not the last one emits a postamble, parking the engine
+ * on its successor's semaphore and publishing that successor's ring tail.
+ *
+ * Return: True if @job emits a ULLS postamble, False otherwise.
+ */
+static inline bool xe_sched_job_ulls_parks(struct xe_sched_job *job)
+{
+ return job->ulls == ULLS_ENTER || job->ulls == ULLS_ACTIVE;
+}
+
+/**
+ * xe_sched_job_ulls_is_chained() - Has a job's predecessor already published it
+ * @job: Xe schedule job object
+ *
+ * A ULLS job which is not the first one has had its ring tail published by its
+ * predecessor's postamble, which also left the engine parked on this job's
+ * semaphore. Submitting it is a semaphore write alone - no H2G and no ring
+ * tail write.
+ *
+ * Return: True if @job was published by its predecessor, False otherwise.
+ */
+static inline bool xe_sched_job_ulls_is_chained(struct xe_sched_job *job)
+{
+ return job->ulls == ULLS_ACTIVE || job->ulls == ULLS_EXIT;
+}
+
bool xe_sched_job_is_migration(struct xe_exec_queue *q);
struct xe_sched_job_snapshot *xe_sched_job_snapshot_capture(struct xe_sched_job *job);
diff --git a/drivers/gpu/drm/xe/xe_sched_job_types.h b/drivers/gpu/drm/xe/xe_sched_job_types.h
index 9f527ac6df3e..c8cdbe68843f 100644
--- a/drivers/gpu/drm/xe/xe_sched_job_types.h
+++ b/drivers/gpu/drm/xe/xe_sched_job_types.h
@@ -49,6 +49,23 @@ struct xe_job_ptrs {
u32 head;
};
+/**
+ * enum xe_ulls_state - ULLS state of a migration job
+ *
+ * Describes where a job sits in a ULLS (Ultra Low Latency Submission)
+ * sequence. See the ULLS documentation in xe_migrate.c.
+ */
+enum xe_ulls_state {
+ /** @ULLS_NONE: Not a ULLS job */
+ ULLS_NONE = 0,
+ /** @ULLS_ENTER: Job which enters ULLS mode */
+ ULLS_ENTER,
+ /** @ULLS_ACTIVE: Job submitted while in ULLS mode */
+ ULLS_ACTIVE,
+ /** @ULLS_EXIT: Job which exits ULLS mode */
+ ULLS_EXIT,
+};
+
/**
* struct xe_sched_job - Xe schedule job (batch buffer tracking)
*/
@@ -79,6 +96,8 @@ struct xe_sched_job {
u32 migrate_flush_flags;
/** @sample_timestamp: Sampling of job timestamp in TDR */
u64 sample_timestamp;
+ /** @ulls: ULLS state of this job */
+ enum xe_ulls_state ulls;
/** @ring_ops_flush_tlb: The ring ops need to flush TLB before payload. */
bool ring_ops_flush_tlb;
/** @ring_ops_force_reset: The ring ops need to trigger a reset before payload. */
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index f450e1c6f750..f66bd1de2b70 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -149,7 +149,7 @@ static bool xe_vm_is_idle(struct xe_vm *vm)
xe_vm_assert_held(vm);
list_for_each_entry(q, &vm->preempt.exec_queues, lr.link) {
- if (!xe_exec_queue_is_idle(q))
+ if (!xe_exec_queue_is_idle(q, 0))
return false;
}
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer
2026-09-25 4:53 ` [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
@ 2026-09-25 6:34 ` sashiko-bot
2026-09-25 7:17 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 6:34 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] drm/xe: Missing ULLS_EXIT in teardown leaves hardware parked on freed memory
--
commit 81f7dc64bc0d604133ac8c3a1f3ec0fd0a9dd8a1
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: Add ULLS migration job support to migration layer
Add function to enter ULLS mode for migration job and delayed worker to
exit (power saving). ULLS mode expected to entered upon page fault or
SVM prefetch. ULLS mode exit delay is currently set to 5ms.
> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> index 471ae5741836e..d7d13a25cdb91 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
[ ... ]
> @@ -95,9 +105,30 @@ struct xe_migrate {
> */
> #define MAX_PTE_PER_SDI 0x1FEU
>
> +static bool xe_migrate_ulls_enabled(struct xe_migrate *m)
> +{
> + lockdep_assert_held(&m->job_mutex);
> + return m->ulls.enabled;
> +}
> +
> +static void xe_migrate_ulls_toggle_enable(struct xe_migrate *m, bool enabled)
> +{
> + lockdep_assert_held(&m->job_mutex);
> + m->ulls.enabled = enabled;
> +}
> +
> static void xe_migrate_fini(void *arg)
> {
> struct xe_migrate *m = arg;
> + struct xe_device *xe = tile_to_xe(m->tile);
> +
> + disable_delayed_work_sync(&m->ulls.exit_work);
> + scoped_guard(mutex, &m->job_mutex) {
> + if (xe_migrate_ulls_enabled(m)) {
> + xe_pm_runtime_put(xe);
> + xe_migrate_ulls_toggle_enable(m, false);
[Severity: High]
Does this missing ULLS_EXIT in the teardown sequence leave the hardware
parked on freed memory?
During driver unbind or probe failure in xe_migrate_fini(), if ULLS mode is
currently active (for example, within 5ms of a recent migration job), the
reserved hardware engine is actively spinning in a MI_SEMAPHORE_WAIT loop.
This teardown path cancels exit_work and clears the ULLS enabled flag, but
fails to push the ULLS_EXIT job or explicitly kill the queue. It then
proceeds to call xe_exec_queue_put() and free the page tables. Because
xe_exec_queue_kill() is never called, GuC is never instructed to preempt
and deregister the context.
Can this regression leave the hardware engine actively executing from the
ring buffer memory while the host frees it, potentially leading to IOMMU CAT
faults, GPU hangs, or host memory corruption?
> + }
> + }
>
> xe_vm_lock(m->q->vm, false);
> xe_bo_unpin(m->pt_bo);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=19
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer
2026-09-25 6:34 ` sashiko-bot
@ 2026-09-25 7:17 ` Matthew Brost
2026-09-25 20:10 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 7:17 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 06:34:11AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] drm/xe: Missing ULLS_EXIT in teardown leaves hardware parked on freed memory
> --
>
> commit 81f7dc64bc0d604133ac8c3a1f3ec0fd0a9dd8a1
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Add ULLS migration job support to migration layer
>
> Add function to enter ULLS mode for migration job and delayed worker to
> exit (power saving). ULLS mode expected to entered upon page fault or
> SVM prefetch. ULLS mode exit delay is currently set to 5ms.
>
> > diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> > index 471ae5741836e..d7d13a25cdb91 100644
> > --- a/drivers/gpu/drm/xe/xe_migrate.c
> > +++ b/drivers/gpu/drm/xe/xe_migrate.c
> [ ... ]
> > @@ -95,9 +105,30 @@ struct xe_migrate {
> > */
> > #define MAX_PTE_PER_SDI 0x1FEU
> >
> > +static bool xe_migrate_ulls_enabled(struct xe_migrate *m)
> > +{
> > + lockdep_assert_held(&m->job_mutex);
> > + return m->ulls.enabled;
> > +}
> > +
> > +static void xe_migrate_ulls_toggle_enable(struct xe_migrate *m, bool enabled)
> > +{
> > + lockdep_assert_held(&m->job_mutex);
> > + m->ulls.enabled = enabled;
> > +}
> > +
> > static void xe_migrate_fini(void *arg)
> > {
> > struct xe_migrate *m = arg;
> > + struct xe_device *xe = tile_to_xe(m->tile);
> > +
> > + disable_delayed_work_sync(&m->ulls.exit_work);
> > + scoped_guard(mutex, &m->job_mutex) {
> > + if (xe_migrate_ulls_enabled(m)) {
> > + xe_pm_runtime_put(xe);
> > + xe_migrate_ulls_toggle_enable(m, false);
>
> [Severity: High]
> Does this missing ULLS_EXIT in the teardown sequence leave the hardware
> parked on freed memory?
>
> During driver unbind or probe failure in xe_migrate_fini(), if ULLS mode is
> currently active (for example, within 5ms of a recent migration job), the
> reserved hardware engine is actively spinning in a MI_SEMAPHORE_WAIT loop.
>
> This teardown path cancels exit_work and clears the ULLS enabled flag, but
> fails to push the ULLS_EXIT job or explicitly kill the queue. It then
> proceeds to call xe_exec_queue_put() and free the page tables. Because
> xe_exec_queue_kill() is never called, GuC is never instructed to preempt
> and deregister the context.
>
> Can this regression leave the hardware engine actively executing from the
> ring buffer memory while the host frees it, potentially leading to IOMMU CAT
> faults, GPU hangs, or host memory corruption?
>
Hmm, we probably should kill the queue or submit a exit job. Let me fix
this.
Matt
> > + }
> > + }
> >
> > xe_vm_lock(m->q->vm, false);
> > xe_bo_unpin(m->pt_bo);
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=19
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer
2026-09-25 7:17 ` Matthew Brost
@ 2026-09-25 20:10 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 20:10 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 12:17:12AM -0700, Matthew Brost wrote:
> On Fri, Sep 25, 2026 at 06:34:11AM +0000, sashiko-bot@kernel.org wrote:
> > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > - [High] drm/xe: Missing ULLS_EXIT in teardown leaves hardware parked on freed memory
> > --
> >
> > commit 81f7dc64bc0d604133ac8c3a1f3ec0fd0a9dd8a1
> > Author: Matthew Brost <matthew.brost@intel.com>
> >
> > drm/xe: Add ULLS migration job support to migration layer
> >
> > Add function to enter ULLS mode for migration job and delayed worker to
> > exit (power saving). ULLS mode expected to entered upon page fault or
> > SVM prefetch. ULLS mode exit delay is currently set to 5ms.
> >
> > > diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> > > index 471ae5741836e..d7d13a25cdb91 100644
> > > --- a/drivers/gpu/drm/xe/xe_migrate.c
> > > +++ b/drivers/gpu/drm/xe/xe_migrate.c
> > [ ... ]
> > > @@ -95,9 +105,30 @@ struct xe_migrate {
> > > */
> > > #define MAX_PTE_PER_SDI 0x1FEU
> > >
> > > +static bool xe_migrate_ulls_enabled(struct xe_migrate *m)
> > > +{
> > > + lockdep_assert_held(&m->job_mutex);
> > > + return m->ulls.enabled;
> > > +}
> > > +
> > > +static void xe_migrate_ulls_toggle_enable(struct xe_migrate *m, bool enabled)
> > > +{
> > > + lockdep_assert_held(&m->job_mutex);
> > > + m->ulls.enabled = enabled;
> > > +}
> > > +
> > > static void xe_migrate_fini(void *arg)
> > > {
> > > struct xe_migrate *m = arg;
> > > + struct xe_device *xe = tile_to_xe(m->tile);
> > > +
> > > + disable_delayed_work_sync(&m->ulls.exit_work);
> > > + scoped_guard(mutex, &m->job_mutex) {
> > > + if (xe_migrate_ulls_enabled(m)) {
> > > + xe_pm_runtime_put(xe);
> > > + xe_migrate_ulls_toggle_enable(m, false);
> >
> > [Severity: High]
> > Does this missing ULLS_EXIT in the teardown sequence leave the hardware
> > parked on freed memory?
> >
> > During driver unbind or probe failure in xe_migrate_fini(), if ULLS mode is
> > currently active (for example, within 5ms of a recent migration job), the
> > reserved hardware engine is actively spinning in a MI_SEMAPHORE_WAIT loop.
> >
> > This teardown path cancels exit_work and clears the ULLS enabled flag, but
> > fails to push the ULLS_EXIT job or explicitly kill the queue. It then
> > proceeds to call xe_exec_queue_put() and free the page tables. Because
> > xe_exec_queue_kill() is never called, GuC is never instructed to preempt
> > and deregister the context.
> >
> > Can this regression leave the hardware engine actively executing from the
> > ring buffer memory while the host frees it, potentially leading to IOMMU CAT
> > faults, GPU hangs, or host memory corruption?
> >
>
> Hmm, we probably should kill the queue or submit a exit job. Let me fix
> this.
On 2nd thought this is actually fine. We put the exec below which
triggers cleanup of the queue. If GuC firmware is running we disable
scheduling and teardown the queue, if the GuC firmware is taken offline
we skip that step, so this this actually all fine. We just need to
cleanup extra runtime PM which this code does.
Matt
>
> Matt
>
> > > + }
> > > + }
> > >
> > > xe_vm_lock(m->q->vm, false);
> > > xe_bo_unpin(m->pt_bo);
> >
> > --
> > Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=19
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (18 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 6:38 ` sashiko-bot
2026-09-25 4:53 ` [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
` (8 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Shuicheng Lin, Himal Prasad Ghimiray
Add preamble and postamble for ULLS migrations jobs. Preamble clears
current semaphore for reuse. Postamble waits on next semaphore which is
set upon next job submission, then advances the ring tail over that job
with an LRI to RING_TAIL, so submitting it costs the CPU nothing beyond
signalling the semaphore.
A job updates the tail on behalf of a successor which has not been
emitted yet, so it cannot know how much ring that successor will occupy.
Pad every ULLS job out to a fixed ULLS_JOB_SIZE_BYTES, which makes the
next tail derivable from where the current job starts. The pad also
supplies the NOPs which must follow an in-ring tail update.
The last ULLS migration job skips BB submission, the postamble and the
tail update (clear current semaphore, write seqno, exit ULLS), padding
the difference so that it still fills a job slot.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
drivers/gpu/drm/xe/xe_ring_ops.c | 75 ++++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_ring_ops_types.h | 24 +++++++++
2 files changed, 99 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_ring_ops.c b/drivers/gpu/drm/xe/xe_ring_ops.c
index 46ab1f0f3564..9f462bfecb83 100644
--- a/drivers/gpu/drm/xe/xe_ring_ops.c
+++ b/drivers/gpu/drm/xe/xe_ring_ops.c
@@ -505,6 +505,68 @@ static void __emit_job_gen12_render_compute(struct xe_sched_job *job,
xe_lrc_write_ring(lrc, dw, i * sizeof(*dw));
}
+static int emit_ulls_preamble(struct xe_lrc *lrc, u32 *dw, int i, u32 seqno)
+{
+ u32 addr = xe_lrc_ulls_semaphore_ggtt_addr(lrc, seqno);
+
+ return emit_store_imm_ggtt(addr, LRC_MIGRATION_ULLS_SEMAPHORE_CLEAR,
+ dw, i);
+}
+
+/*
+ * Advance the ring tail from within the ring, so submitting the next ULLS job
+ * needs nothing from the CPU beyond signalling the semaphore. All ULLS jobs
+ * occupy exactly ULLS_JOB_SIZE_BYTES, so the tail the next job ends at is two
+ * job slots on from where this job started, even though that job has not been
+ * emitted yet. Both LRC and MMIO ring tail advanced in step.
+ */
+static int emit_ulls_ring_tail(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
+ int i, u32 head)
+{
+ u32 next_tail = (head + 2 * ULLS_JOB_SIZE_BYTES) & (lrc->ring.size - 1);
+
+ xe_gt_assert(gt, IS_ALIGNED(next_tail, 8));
+
+ i = emit_store_imm_ggtt(xe_lrc_ring_tail_ggtt_addr(lrc), next_tail,
+ dw, i);
+
+ dw[i++] = MI_LOAD_REGISTER_IMM | MI_LRI_NUM_REGS(1) |
+ MI_LRI_LRM_CS_MMIO;
+ dw[i++] = RING_TAIL(0).addr;
+ dw[i++] = next_tail;
+
+ return i;
+}
+
+/* Publish the next job's tail, then park the engine on its semaphore */
+static int emit_ulls_postamble(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
+ int i, u32 seqno, u32 head)
+{
+ i = emit_ulls_ring_tail(gt, lrc, dw, i, head);
+
+ dw[i++] = MI_SEMAPHORE_WAIT |
+ MI_SEMW_GGTT |
+ MI_SEMW_POLL |
+ MI_SEMW_COMPARE(SAD_EQ_SDD);
+ dw[i++] = LRC_MIGRATION_ULLS_SEMAPHORE_SIGNAL;
+ dw[i++] = xe_lrc_ulls_semaphore_ggtt_addr(lrc, seqno + 1);
+ dw[i++] = 0;
+ dw[i++] = 0;
+
+ return i;
+}
+
+/* Pad out to the fixed ULLS job size */
+static int emit_ulls_pad(struct xe_gt *gt, u32 *dw, int i)
+{
+ xe_gt_assert(gt, i <= ULLS_JOB_SIZE_DW);
+
+ while (i < ULLS_JOB_SIZE_DW)
+ dw[i++] = MI_NOOP;
+
+ return i;
+}
+
static void emit_migration_job_gen12(struct xe_sched_job *job,
struct xe_lrc *lrc, u32 *head,
u32 seqno)
@@ -518,10 +580,16 @@ static void emit_migration_job_gen12(struct xe_sched_job *job,
xe_gt_assert(gt, !job->ring_ops_force_reset);
+ if (xe_sched_job_is_ulls(job))
+ i = emit_ulls_preamble(lrc, dw, i, seqno);
+
i = emit_copy_timestamp(xe, lrc, dw, i);
i = emit_store_imm_ggtt(saddr, seqno, dw, i);
+ if (!xe_sched_job_ulls_has_batch(job))
+ goto seqno_write;
+
dw[i++] = MI_ARB_ON_OFF | MI_ARB_DISABLE; /* Enabled again below */
i = emit_bb_start(job->ptrs[0].batch_addr, BIT(8), dw, i);
@@ -532,12 +600,19 @@ static void emit_migration_job_gen12(struct xe_sched_job *job,
i = emit_bb_start(job->ptrs[1].batch_addr, BIT(8), dw, i);
+seqno_write:
i = emit_flush_imm_ggtt(xe_lrc_seqno_ggtt_addr(lrc), seqno,
job->migrate_flush_flags,
dw, i);
i = emit_user_interrupt(dw, i);
+ if (xe_sched_job_ulls_parks(job))
+ i = emit_ulls_postamble(gt, lrc, dw, i, seqno, *head);
+
+ if (xe_sched_job_is_ulls(job))
+ i = emit_ulls_pad(gt, dw, i);
+
xe_gt_assert(job->q->gt, i <= MAX_JOB_SIZE_DW);
xe_lrc_write_ring(lrc, dw, i * sizeof(*dw));
diff --git a/drivers/gpu/drm/xe/xe_ring_ops_types.h b/drivers/gpu/drm/xe/xe_ring_ops_types.h
index 52ff96bc4100..ea4af321dd7c 100644
--- a/drivers/gpu/drm/xe/xe_ring_ops_types.h
+++ b/drivers/gpu/drm/xe/xe_ring_ops_types.h
@@ -6,6 +6,7 @@
#ifndef _XE_RING_OPS_TYPES_H_
#define _XE_RING_OPS_TYPES_H_
+#include <linux/build_bug.h>
#include <linux/types.h>
struct xe_gt;
@@ -14,6 +15,29 @@ struct xe_sched_job;
#define MAX_JOB_SIZE_DW 74
#define MAX_JOB_SIZE_BYTES (MAX_JOB_SIZE_DW * 4)
+/*
+ * ULLS migration jobs advance the ring tail from within the ring itself, so a
+ * job has to know where its successor will end before that successor has been
+ * emitted. Every ULLS job is therefore padded to a fixed size, letting the
+ * next tail be derived arithmetically.
+ *
+ * Sized for the largest such job, emitted by emit_migration_job_gen12():
+ * preamble (4), copy timestamp (8, its size on an SRIOV VF), start seqno
+ * store (4), arbitration off (1), batch buffer starts (2 * 3), pre-parser
+ * bracketed flush invalidate (6), seqno flush (4), user interrupt (3) and
+ * postamble (7 + 5).
+ */
+#define ULLS_JOB_SIZE_DW 48
+#define ULLS_JOB_SIZE_BYTES (ULLS_JOB_SIZE_DW * 4)
+
+/*
+ * RING_TAIL only encodes a qword aligned offset, and xe_lrc_write_ring()
+ * appends a NOP to anything shorter, either of which would desynchronise the
+ * ring from the tail a job predicts for its successor.
+ */
+static_assert(ULLS_JOB_SIZE_BYTES % 8 == 0);
+static_assert(ULLS_JOB_SIZE_DW <= MAX_JOB_SIZE_DW);
+
/**
* struct xe_ring_ops - Ring operations
*/
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops
2026-09-25 4:53 ` [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
@ 2026-09-25 6:38 ` sashiko-bot
2026-09-25 7:02 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 6:38 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] drm/xe: GPU hang from CS prefetching uninitialized ULLS jobs
--
commit ca5d0374fab1065ae65164d64f6d5510acadd964
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: Add ULLS migration job support to ring ops
Add preamble and postamble for ULLS migrations jobs. Preamble clears
current semaphore for reuse. Postamble waits on next semaphore which is
set upon next job submission, then advances the ring tail over that job
with an LRI to RING_TAIL, so submitting it costs the CPU nothing beyond
signalling the semaphore.
> diff --git a/drivers/gpu/drm/xe/xe_ring_ops.c b/drivers/gpu/drm/xe/xe_ring_ops.c
> index 46ab1f0f35648..9f462bfecb831 100644
> --- a/drivers/gpu/drm/xe/xe_ring_ops.c
> +++ b/drivers/gpu/drm/xe/xe_ring_ops.c
[ ... ]
> +/* Publish the next job's tail, then park the engine on its semaphore */
> +static int emit_ulls_postamble(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
> + int i, u32 seqno, u32 head)
> +{
> + i = emit_ulls_ring_tail(gt, lrc, dw, i, head);
> +
> + dw[i++] = MI_SEMAPHORE_WAIT |
> + MI_SEMW_GGTT |
> + MI_SEMW_POLL |
> + MI_SEMW_COMPARE(SAD_EQ_SDD);
> + dw[i++] = LRC_MIGRATION_ULLS_SEMAPHORE_SIGNAL;
[Severity: High]
Does updating the ring tail to the next job's boundary before blocking on
the semaphore open a window where the Command Streamer (CS) prefetcher can
fetch uninitialized memory into its instruction cache?
When the CPU later writes the next job and signals the semaphore, the
absence of an MMIO tail update means the CS cache might not be invalidated.
Could this cause the CS to execute stale data and hang the GPU?
[ ... ]
> +/* Pad out to the fixed ULLS job size */
> +static int emit_ulls_pad(struct xe_gt *gt, u32 *dw, int i)
> +{
> + xe_gt_assert(gt, i <= ULLS_JOB_SIZE_DW);
> +
> + while (i < ULLS_JOB_SIZE_DW)
> + dw[i++] = MI_NOOP;
> +
> + return i;
> +}
[Severity: High]
If a ULLS job reaches the maximum size of ULLS_JOB_SIZE_DW (which is 48),
this loop will provide 0 NOOPs of padding.
Will this fail to buffer the prefetcher from the next job's uninitialized
memory, exacerbating the prefetching issue mentioned above?
[ ... ]
> diff --git a/drivers/gpu/drm/xe/xe_ring_ops_types.h b/drivers/gpu/drm/xe/xe_ring_ops_types.h
> index 52ff96bc41004..ea4af321dd7c4 100644
> --- a/drivers/gpu/drm/xe/xe_ring_ops_types.h
> +++ b/drivers/gpu/drm/xe/xe_ring_ops_types.h
[ ... ]
> +/*
> + * ULLS migration jobs advance the ring tail from within the ring itself, so a
> + * job has to know where its successor will end before that successor has been
> + * emitted. Every ULLS job is therefore padded to a fixed size, letting the
> + * next tail be derived arithmetically.
> + *
> + * Sized for the largest such job, emitted by emit_migration_job_gen12():
> + * preamble (4), copy timestamp (8, its size on an SRIOV VF), start seqno
> + * store (4), arbitration off (1), batch buffer starts (2 * 3), pre-parser
> + * bracketed flush invalidate (6), seqno flush (4), user interrupt (3) and
> + * postamble (7 + 5).
> + */
> +#define ULLS_JOB_SIZE_DW 48
> +#define ULLS_JOB_SIZE_BYTES (ULLS_JOB_SIZE_DW * 4)
[Severity: High]
Since this maximum size is exactly equal to the maximal job length, it
guarantees no NOOP padding is appended for the largest jobs.
Should this maximum size account for additional NOOPs to safely pad the
CS prefetch window?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=20
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops
2026-09-25 6:38 ` sashiko-bot
@ 2026-09-25 7:02 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 7:02 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 06:38:44AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] drm/xe: GPU hang from CS prefetching uninitialized ULLS jobs
> --
>
> commit ca5d0374fab1065ae65164d64f6d5510acadd964
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Add ULLS migration job support to ring ops
>
> Add preamble and postamble for ULLS migrations jobs. Preamble clears
> current semaphore for reuse. Postamble waits on next semaphore which is
> set upon next job submission, then advances the ring tail over that job
> with an LRI to RING_TAIL, so submitting it costs the CPU nothing beyond
> signalling the semaphore.
>
> > diff --git a/drivers/gpu/drm/xe/xe_ring_ops.c b/drivers/gpu/drm/xe/xe_ring_ops.c
> > index 46ab1f0f35648..9f462bfecb831 100644
> > --- a/drivers/gpu/drm/xe/xe_ring_ops.c
> > +++ b/drivers/gpu/drm/xe/xe_ring_ops.c
> [ ... ]
> > +/* Publish the next job's tail, then park the engine on its semaphore */
> > +static int emit_ulls_postamble(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
> > + int i, u32 seqno, u32 head)
> > +{
> > + i = emit_ulls_ring_tail(gt, lrc, dw, i, head);
> > +
> > + dw[i++] = MI_SEMAPHORE_WAIT |
> > + MI_SEMW_GGTT |
> > + MI_SEMW_POLL |
> > + MI_SEMW_COMPARE(SAD_EQ_SDD);
> > + dw[i++] = LRC_MIGRATION_ULLS_SEMAPHORE_SIGNAL;
>
> [Severity: High]
All semmly non-issues during testing. I believe the semaphore wait acts
as prefetch flush.
Matt
> Does updating the ring tail to the next job's boundary before blocking on
> the semaphore open a window where the Command Streamer (CS) prefetcher can
> fetch uninitialized memory into its instruction cache?
>
> When the CPU later writes the next job and signals the semaphore, the
> absence of an MMIO tail update means the CS cache might not be invalidated.
> Could this cause the CS to execute stale data and hang the GPU?
>
> [ ... ]
> > +/* Pad out to the fixed ULLS job size */
> > +static int emit_ulls_pad(struct xe_gt *gt, u32 *dw, int i)
> > +{
> > + xe_gt_assert(gt, i <= ULLS_JOB_SIZE_DW);
> > +
> > + while (i < ULLS_JOB_SIZE_DW)
> > + dw[i++] = MI_NOOP;
> > +
> > + return i;
> > +}
>
> [Severity: High]
> If a ULLS job reaches the maximum size of ULLS_JOB_SIZE_DW (which is 48),
> this loop will provide 0 NOOPs of padding.
>
> Will this fail to buffer the prefetcher from the next job's uninitialized
> memory, exacerbating the prefetching issue mentioned above?
>
> [ ... ]
> > diff --git a/drivers/gpu/drm/xe/xe_ring_ops_types.h b/drivers/gpu/drm/xe/xe_ring_ops_types.h
> > index 52ff96bc41004..ea4af321dd7c4 100644
> > --- a/drivers/gpu/drm/xe/xe_ring_ops_types.h
> > +++ b/drivers/gpu/drm/xe/xe_ring_ops_types.h
> [ ... ]
> > +/*
> > + * ULLS migration jobs advance the ring tail from within the ring itself, so a
> > + * job has to know where its successor will end before that successor has been
> > + * emitted. Every ULLS job is therefore padded to a fixed size, letting the
> > + * next tail be derived arithmetically.
> > + *
> > + * Sized for the largest such job, emitted by emit_migration_job_gen12():
> > + * preamble (4), copy timestamp (8, its size on an SRIOV VF), start seqno
> > + * store (4), arbitration off (1), batch buffer starts (2 * 3), pre-parser
> > + * bracketed flush invalidate (6), seqno flush (4), user interrupt (3) and
> > + * postamble (7 + 5).
> > + */
> > +#define ULLS_JOB_SIZE_DW 48
> > +#define ULLS_JOB_SIZE_BYTES (ULLS_JOB_SIZE_DW * 4)
>
> [Severity: High]
> Since this maximum size is exactly equal to the maximal job length, it
> guarantees no NOOP padding is appended for the largest jobs.
>
> Should this maximum size account for additional NOOPs to safely pad the
> CS prefetch window?
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=20
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (19 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 6:48 ` sashiko-bot
2026-09-25 4:53 ` [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
` (7 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Himal Prasad Ghimiray, Shuicheng Lin
Add ULLS migration job support to GuC submission backend.
Changes required:
- On migration queue, reduce max jobs to the number of ULLS semaphores
minus one
- Skip writing the saved ring tail for ULLS jobs except for the first
ULLS job - the ring updates both the saved tail and the tail register
itself
- Set ULLS sempahore for current job releasing last job except for first
ULLS job
- Suppress submit H2G for ULLS except for first ULLS job
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
---
drivers/gpu/drm/xe/xe_guc_submit.c | 17 +++++++++++++----
1 file changed, 13 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index a2784763e72b..cd9bb9f9e28c 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -1190,9 +1190,10 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
xe_gt_assert(guc_to_gt(guc), exec_queue_registered(q));
if (!job->restore_replay || job->last_replay) {
+ /* A ULLS job past the first publishes its own ring tail */
if (xe_exec_queue_is_parallel(q))
wq_item_append(q);
- else
+ else if (!xe_sched_job_ulls_is_chained(job))
xe_lrc_set_ring_tail(lrc, lrc->ring.tail);
job->last_replay = false;
}
@@ -1209,6 +1210,9 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
if (exec_queue_suspended(q))
return;
+ if (xe_sched_job_ulls_is_chained(job))
+ xe_lrc_set_ulls_semaphore(lrc, xe_sched_job_lrc_seqno(job));
+
if (!exec_queue_enabled(q)) {
action[len++] = XE_GUC_ACTION_SCHED_CONTEXT_MODE_SET;
action[len++] = q->guc->id;
@@ -1222,13 +1226,14 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
set_exec_queue_pending_enable(q);
set_exec_queue_enabled(q);
trace_xe_exec_queue_scheduling_enable(q);
- } else {
+ } else if (!xe_sched_job_ulls_is_chained(job)) {
action[len++] = XE_GUC_ACTION_SCHED_CONTEXT;
action[len++] = q->guc->id;
trace_xe_exec_queue_submit(q);
}
- xe_guc_ct_send(&guc->ct, action, len, g2h_len, num_g2h);
+ if (!xe_sched_job_ulls_is_chained(job) || num_g2h)
+ xe_guc_ct_send(&guc->ct, action, len, g2h_len, num_g2h);
if (extra_submit) {
len = 0;
@@ -2100,6 +2105,7 @@ static int guc_exec_queue_init(struct xe_exec_queue *q)
struct xe_guc_exec_queue *ge;
long timeout;
int err, i;
+ int max_jobs = (xe_lrc_ring_size() / MAX_JOB_SIZE_BYTES);
xe_gt_assert(guc_to_gt(guc), xe_device_uc_enabled(guc_to_xe(guc)));
@@ -2139,8 +2145,11 @@ static int guc_exec_queue_init(struct xe_exec_queue *q)
submit_wq = primary->guc->sched.base.submit_wq;
}
+ if (q->vm && q->vm->flags & XE_VM_FLAG_MIGRATION)
+ max_jobs = min(max_jobs, LRC_MIGRATION_ULLS_SEMAPHORE_COUNT - 1);
+
err = xe_sched_init(&ge->sched, &drm_sched_ops, &xe_sched_ops,
- submit_wq, xe_lrc_ring_size() / MAX_JOB_SIZE_BYTES, 64,
+ submit_wq, max_jobs, 64,
timeout, guc_to_gt(guc)->ordered_wq, NULL,
ge->name, gt_to_xe(q->gt)->drm.dev);
if (err)
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission
2026-09-25 4:53 ` [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
@ 2026-09-25 6:48 ` sashiko-bot
2026-09-25 7:08 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 6:48 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] drm/xe: TOCTOU race in ULLS migration job prefetching leads to stale command execution
--
commit 231db353d11582f236795322ee38df1ad9cad0d1
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: Add ULLS migration job support to GuC submission
This commit adds Ultra Low Latency Submission (ULLS) migration job support to
the GuC submission backend. It introduces optimizations for chained ULLS jobs
by having the GPU publish its own ring tail and suppressing the CPU's RING_TAIL
MMIO updates.
> diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> index a2784763e72ba..cd9bb9f9e28c8 100644
> --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
> @@ -1190,9 +1190,10 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
> xe_gt_assert(guc_to_gt(guc), exec_queue_registered(q));
>
> if (!job->restore_replay || job->last_replay) {
> + /* A ULLS job past the first publishes its own ring tail */
> if (xe_exec_queue_is_parallel(q))
> wq_item_append(q);
> - else
> + else if (!xe_sched_job_ulls_is_chained(job))
> xe_lrc_set_ring_tail(lrc, lrc->ring.tail);
[Severity: High]
Does skipping the RING_TAIL MMIO write here expose a race condition where the
GPU could execute stale ring buffer memory?
For ULLS chained jobs, the GPU advances the tail via MI_LRI prior to executing
MI_SEMAPHORE_WAIT, which happens in emit_ulls_postamble():
drivers/gpu/drm/xe/xe_ring_ops.c:emit_ulls_postamble()
i = emit_ulls_ring_tail(gt, lrc, dw, i, head);
dw[i++] = MI_SEMAPHORE_WAIT | ...
Because the CPU intentionally skips the xe_lrc_set_ring_tail() MMIO write for
these chained jobs, the hardware mechanism that normally invalidates the GPU's
instruction cache is bypassed.
Could the Command Streamer fetch unit see the advanced RING_TAIL and prefetch
memory for the next job before the CPU has actually written it to the ring
buffer? If the CPU hasn't populated the commands yet, when the CPU finally
writes the next job and signals the semaphore, it appears the GPU might wake up
and execute the stale prefetched instructions from its internal cache rather
than the newly written migration commands.
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=21
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission
2026-09-25 6:48 ` sashiko-bot
@ 2026-09-25 7:08 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 7:08 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 06:48:57AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] drm/xe: TOCTOU race in ULLS migration job prefetching leads to stale command execution
> --
>
> commit 231db353d11582f236795322ee38df1ad9cad0d1
> Author: Matthew Brost <matthew.brost@intel.com>
> drm/xe: Add ULLS migration job support to GuC submission
>
> This commit adds Ultra Low Latency Submission (ULLS) migration job support to
> the GuC submission backend. It introduces optimizations for chained ULLS jobs
> by having the GPU publish its own ring tail and suppressing the CPU's RING_TAIL
> MMIO updates.
>
> > diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> > index a2784763e72ba..cd9bb9f9e28c8 100644
> > --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> > +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
> > @@ -1190,9 +1190,10 @@ static void submit_exec_queue(struct xe_exec_queue *q, struct xe_sched_job *job)
> > xe_gt_assert(guc_to_gt(guc), exec_queue_registered(q));
> >
> > if (!job->restore_replay || job->last_replay) {
> > + /* A ULLS job past the first publishes its own ring tail */
> > if (xe_exec_queue_is_parallel(q))
> > wq_item_append(q);
> > - else
> > + else if (!xe_sched_job_ulls_is_chained(job))
> > xe_lrc_set_ring_tail(lrc, lrc->ring.tail);
>
> [Severity: High]
Not observed in testing - semaphore wait acts as prefetch barrier.
Before semaphore release all CPU writes are GPU visable.
Matt
> Does skipping the RING_TAIL MMIO write here expose a race condition where the
> GPU could execute stale ring buffer memory?
>
> For ULLS chained jobs, the GPU advances the tail via MI_LRI prior to executing
> MI_SEMAPHORE_WAIT, which happens in emit_ulls_postamble():
>
> drivers/gpu/drm/xe/xe_ring_ops.c:emit_ulls_postamble()
> i = emit_ulls_ring_tail(gt, lrc, dw, i, head);
> dw[i++] = MI_SEMAPHORE_WAIT | ...
>
> Because the CPU intentionally skips the xe_lrc_set_ring_tail() MMIO write for
> these chained jobs, the hardware mechanism that normally invalidates the GPU's
> instruction cache is bypassed.
>
> Could the Command Streamer fetch unit see the advanced RING_TAIL and prefetch
> memory for the next job before the CPU has actually written it to the ring
> buffer? If the CPU hasn't populated the commands yet, when the CPU finally
> writes the next job and signals the semaphore, it appears the GPU might wake up
> and execute the stale prefetched instructions from its internal cache rather
> than the newly written migration commands.
>
> [ ... ]
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=21
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (20 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 17:49 ` Maarten Lankhorst
2026-09-25 4:53 ` [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute Matthew Brost
` (6 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe; +Cc: Shuicheng Lin, Himal Prasad Ghimiray
Call xe_migrate_ulls_enter upon page fault or SVM prefetch in an
effort speed up these critical paths.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
---
v7:
- s/migration/migrate (Shuicheng)
---
drivers/gpu/drm/xe/xe_pagefault.c | 3 +++
drivers/gpu/drm/xe/xe_vm.c | 4 +++-
2 files changed, 6 insertions(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
index f9e5bd41cfd1..f219e9e3b73e 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_pagefault.c
@@ -15,6 +15,7 @@
#include "xe_gt_stats.h"
#include "xe_hw_engine.h"
#include "xe_log.h"
+#include "xe_migrate.h"
#include "xe_pagefault.h"
#include "xe_pagefault_types.h"
#include "xe_pm.h"
@@ -281,6 +282,8 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
if (IS_ERR(vm))
return PTR_ERR(vm);
+ xe_migrate_ulls_enter(gt_to_tile(gt)->migrate);
+
down_read(&vm->lock);
if (xe_vm_is_closed(vm)) {
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index f66bd1de2b70..425c678f4480 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -2528,8 +2528,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
ctx.devmem_possible = IS_DGFX(vm->xe) &&
IS_ENABLED(CONFIG_DRM_XE_PAGEMAP);
- for_each_tile(tile, vm->xe, id)
+ for_each_tile(tile, vm->xe, id) {
+ xe_migrate_ulls_enter(tile->migrate);
tile_mask |= 0x1 << id;
+ }
if (prefetch_region == DRM_XE_CONSULT_MEM_ADVISE_PREF_LOC) {
dpagemap = xe_vma_resolve_pagemap(vma,
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
2026-09-25 4:53 ` [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
@ 2026-09-25 17:49 ` Maarten Lankhorst
2026-09-25 18:17 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: Maarten Lankhorst @ 2026-09-25 17:49 UTC (permalink / raw)
To: Matthew Brost, intel-xe; +Cc: Shuicheng Lin, Himal Prasad Ghimiray
Hey,
I know this has been reviewed, but can it be changed to have xe_migrate_ulls_enter + xe_migrate_ulls_leave?
It decreases the amount of pingponging with debugging enabled.
I kept below patch in my tree that did just that, feel free to merge with this patch.
Kind regards,
~Maarten Lankhorst
commit 6beeb2f7ff909bdaf9f8b1eeaa014cd15a65c78b
Author: Maarten Lankhorst <dev@lankhorst.se>
Date: Mon Jun 8 19:05:46 2026 +0200
drm/xe: Keep ulls alive while servicing ops.
Instead of bumping at the start, when performing a lot of work ulls may
run longer, so keep track using a refcount to keep the full benefit of
ulls.
Signed-off-by: Maarten Lankhorst <dev@lankhorst.se>
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index d28da51d71452..f0b6641dd9815 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -58,6 +58,8 @@ struct xe_migrate {
struct xe_tile *tile;
/** @job_mutex: Timeline mutex for @eng. */
struct mutex job_mutex;
+ /** @ulls_used: How many outstanding ulls operations there are */
+ atomic_t ulls_used;
/** @pt_bo: Page-table buffer object. */
struct xe_bo *pt_bo;
/** @batch_base_ofs: VM offset of the migration batch buffer */
@@ -475,6 +477,8 @@ static int xe_migrate_lock_prepare_vm(struct xe_tile *tile, struct xe_migrate *m
*
* If DGFX and not a VF, enter ULLS mode bypassing GuC / HW context
* switches by utilizing semaphore and continuously running batches.
+ *
+ * Pairs with xe_migrate_ulls_leave().
*/
void xe_migrate_ulls_enter(struct xe_migrate *m)
{
@@ -488,6 +492,12 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
return;
+ /* Bump */
+ if (atomic_inc_return(&m->ulls_used) > 1)
+ return;
+
+ cancel_delayed_work_sync(&m->ulls.exit_work);
+
job_alloc:
if (alloc) {
/*
@@ -534,10 +544,30 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
}
if (job)
xe_sched_job_put(job);
+ mutex_unlock(&m->job_mutex);
+}
+
+/**
+ * xe_migrate_ulls_leave() - Leave ULLS mode
+ * @m: The migration context.
+ *
+ * Leaves the critical part of the migration context, it may continue to be enabled.
+ * Pairs with xe_migrate_ulls_enter().
+ */
+void xe_migrate_ulls_leave(struct xe_migrate *m)
+{
+ struct xe_device *xe = tile_to_xe(m->tile);
+
+ if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
+ return;
+
+ if (atomic_dec_return(&m->ulls_used))
+ return;
+
+ guard(mutex)(&m->job_mutex);
if (m->ulls.enabled)
mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
ULLS_EXIT_JIFFIES);
- mutex_unlock(&m->job_mutex);
}
static void xe_migrate_ulls_exit(struct work_struct *work)
@@ -599,8 +629,8 @@ static void xe_migrate_ulls_exit(struct work_struct *work)
ULLS_EXIT_JIFFIES);
}
- drm_dev_exit(idx);
mutex_unlock(&m->job_mutex);
+ drm_dev_exit(idx);
}
/**
diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
index 2e8be10fdb71f..c909e1afb476e 100644
--- a/drivers/gpu/drm/xe/xe_migrate.h
+++ b/drivers/gpu/drm/xe/xe_migrate.h
@@ -98,5 +98,6 @@ int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
#endif
void xe_migrate_ulls_enter(struct xe_migrate *m);
+void xe_migrate_ulls_leave(struct xe_migrate *m);
#endif
diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
index 782ed50a2a43e..51b2cb8fdcf1c 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_pagefault.c
@@ -315,6 +315,7 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
unlock_vm:
up_read(&vm->lock);
+ xe_migrate_ulls_leave(gt_to_tile(gt)->migrate);
xe_vm_put(vm);
return err;
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index a1eca4871175b..fc9f587a35905 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -2557,7 +2557,7 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
if (addr)
goto alloc_next_range;
else
- goto print_op_label;
+ goto ulls_leave;
}
if (IS_ERR(svm_range)) {
@@ -2598,8 +2598,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
if (need_put)
xe_svm_range_put(svm_range);
+ulls_leave:
+ for_each_tile(tile, vm->xe, id)
+ xe_migrate_ulls_leave(tile->migrate);
}
-print_op_label:
print_op(vm->xe, __op);
}
On 9/25/26 06:53, Matthew Brost wrote:
> Call xe_migrate_ulls_enter upon page fault or SVM prefetch in an
> effort speed up these critical paths.
>
> Signed-off-by: Matthew Brost <matthew.brost@intel.com>
> Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
>
> ---
> v7:
> - s/migration/migrate (Shuicheng)
> ---
> drivers/gpu/drm/xe/xe_pagefault.c | 3 +++
> drivers/gpu/drm/xe/xe_vm.c | 4 +++-
> 2 files changed, 6 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
> index f9e5bd41cfd1..f219e9e3b73e 100644
> --- a/drivers/gpu/drm/xe/xe_pagefault.c
> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
> @@ -15,6 +15,7 @@
> #include "xe_gt_stats.h"
> #include "xe_hw_engine.h"
> #include "xe_log.h"
> +#include "xe_migrate.h"
> #include "xe_pagefault.h"
> #include "xe_pagefault_types.h"
> #include "xe_pm.h"
> @@ -281,6 +282,8 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
> if (IS_ERR(vm))
> return PTR_ERR(vm);
>
> + xe_migrate_ulls_enter(gt_to_tile(gt)->migrate);
> +
> down_read(&vm->lock);
>
> if (xe_vm_is_closed(vm)) {
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index f66bd1de2b70..425c678f4480 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
> @@ -2528,8 +2528,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> ctx.devmem_possible = IS_DGFX(vm->xe) &&
> IS_ENABLED(CONFIG_DRM_XE_PAGEMAP);
>
> - for_each_tile(tile, vm->xe, id)
> + for_each_tile(tile, vm->xe, id) {
> + xe_migrate_ulls_enter(tile->migrate);
> tile_mask |= 0x1 << id;
> + }
>
> if (prefetch_region == DRM_XE_CONSULT_MEM_ADVISE_PREF_LOC) {
> dpagemap = xe_vma_resolve_pagemap(vma,
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
2026-09-25 17:49 ` Maarten Lankhorst
@ 2026-09-25 18:17 ` Matthew Brost
2026-09-25 18:26 ` Maarten Lankhorst
0 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 18:17 UTC (permalink / raw)
To: Maarten Lankhorst; +Cc: intel-xe, Shuicheng Lin, Himal Prasad Ghimiray
On Fri, Sep 25, 2026 at 07:49:43PM +0200, Maarten Lankhorst wrote:
> Hey,
>
> I know this has been reviewed, but can it be changed to have xe_migrate_ulls_enter + xe_migrate_ulls_leave?
> It decreases the amount of pingponging with debugging enabled.
>
> I kept below patch in my tree that did just that, feel free to merge with this patch.
>
Ok, I see the potential races where before we get to a copy job in
either prefetch or a pagefault before the exit period ULLS will disarm
itself - I'd hope within the default of 5ms that would never happen in
practice as 5ms is a *really* long time in hot paths (prefetch, SVM).
Is this the issue you are seeing in debug builds, I personally have
never noticed this but generally run with perf builds.
Are you using ULLS somewhere else where 5ms can be hit more easier?
Want to get a full picture before commiting to something.
Matt
> Kind regards,
> ~Maarten Lankhorst
>
> commit 6beeb2f7ff909bdaf9f8b1eeaa014cd15a65c78b
> Author: Maarten Lankhorst <dev@lankhorst.se>
> Date: Mon Jun 8 19:05:46 2026 +0200
>
> drm/xe: Keep ulls alive while servicing ops.
>
> Instead of bumping at the start, when performing a lot of work ulls may
> run longer, so keep track using a refcount to keep the full benefit of
> ulls.
>
> Signed-off-by: Maarten Lankhorst <dev@lankhorst.se>
>
> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> index d28da51d71452..f0b6641dd9815 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
> @@ -58,6 +58,8 @@ struct xe_migrate {
> struct xe_tile *tile;
> /** @job_mutex: Timeline mutex for @eng. */
> struct mutex job_mutex;
> + /** @ulls_used: How many outstanding ulls operations there are */
> + atomic_t ulls_used;
> /** @pt_bo: Page-table buffer object. */
> struct xe_bo *pt_bo;
> /** @batch_base_ofs: VM offset of the migration batch buffer */
> @@ -475,6 +477,8 @@ static int xe_migrate_lock_prepare_vm(struct xe_tile *tile, struct xe_migrate *m
> *
> * If DGFX and not a VF, enter ULLS mode bypassing GuC / HW context
> * switches by utilizing semaphore and continuously running batches.
> + *
> + * Pairs with xe_migrate_ulls_leave().
> */
> void xe_migrate_ulls_enter(struct xe_migrate *m)
> {
> @@ -488,6 +492,12 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
> if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
> return;
>
> + /* Bump */
> + if (atomic_inc_return(&m->ulls_used) > 1)
> + return;
> +
> + cancel_delayed_work_sync(&m->ulls.exit_work);
> +
> job_alloc:
> if (alloc) {
> /*
> @@ -534,10 +544,30 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
> }
> if (job)
> xe_sched_job_put(job);
> + mutex_unlock(&m->job_mutex);
> +}
> +
> +/**
> + * xe_migrate_ulls_leave() - Leave ULLS mode
> + * @m: The migration context.
> + *
> + * Leaves the critical part of the migration context, it may continue to be enabled.
> + * Pairs with xe_migrate_ulls_enter().
> + */
> +void xe_migrate_ulls_leave(struct xe_migrate *m)
> +{
> + struct xe_device *xe = tile_to_xe(m->tile);
> +
> + if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
> + return;
> +
> + if (atomic_dec_return(&m->ulls_used))
> + return;
> +
> + guard(mutex)(&m->job_mutex);
> if (m->ulls.enabled)
> mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
> ULLS_EXIT_JIFFIES);
> - mutex_unlock(&m->job_mutex);
> }
>
> static void xe_migrate_ulls_exit(struct work_struct *work)
> @@ -599,8 +629,8 @@ static void xe_migrate_ulls_exit(struct work_struct *work)
> ULLS_EXIT_JIFFIES);
> }
>
> - drm_dev_exit(idx);
> mutex_unlock(&m->job_mutex);
> + drm_dev_exit(idx);
> }
>
> /**
> diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
> index 2e8be10fdb71f..c909e1afb476e 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.h
> +++ b/drivers/gpu/drm/xe/xe_migrate.h
> @@ -98,5 +98,6 @@ int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
> #endif
>
> void xe_migrate_ulls_enter(struct xe_migrate *m);
> +void xe_migrate_ulls_leave(struct xe_migrate *m);
>
> #endif
> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
> index 782ed50a2a43e..51b2cb8fdcf1c 100644
> --- a/drivers/gpu/drm/xe/xe_pagefault.c
> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
> @@ -315,6 +315,7 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
>
> unlock_vm:
> up_read(&vm->lock);
> + xe_migrate_ulls_leave(gt_to_tile(gt)->migrate);
> xe_vm_put(vm);
>
> return err;
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index a1eca4871175b..fc9f587a35905 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
> @@ -2557,7 +2557,7 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> if (addr)
> goto alloc_next_range;
> else
> - goto print_op_label;
> + goto ulls_leave;
> }
>
> if (IS_ERR(svm_range)) {
> @@ -2598,8 +2598,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> if (need_put)
> xe_svm_range_put(svm_range);
>
> +ulls_leave:
> + for_each_tile(tile, vm->xe, id)
> + xe_migrate_ulls_leave(tile->migrate);
> }
> -print_op_label:
> print_op(vm->xe, __op);
> }
>
>
>
>
> On 9/25/26 06:53, Matthew Brost wrote:
> > Call xe_migrate_ulls_enter upon page fault or SVM prefetch in an
> > effort speed up these critical paths.
> >
> > Signed-off-by: Matthew Brost <matthew.brost@intel.com>
> > Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
> > Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> >
> > ---
> > v7:
> > - s/migration/migrate (Shuicheng)
> > ---
> > drivers/gpu/drm/xe/xe_pagefault.c | 3 +++
> > drivers/gpu/drm/xe/xe_vm.c | 4 +++-
> > 2 files changed, 6 insertions(+), 1 deletion(-)
> >
> > diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
> > index f9e5bd41cfd1..f219e9e3b73e 100644
> > --- a/drivers/gpu/drm/xe/xe_pagefault.c
> > +++ b/drivers/gpu/drm/xe/xe_pagefault.c
> > @@ -15,6 +15,7 @@
> > #include "xe_gt_stats.h"
> > #include "xe_hw_engine.h"
> > #include "xe_log.h"
> > +#include "xe_migrate.h"
> > #include "xe_pagefault.h"
> > #include "xe_pagefault_types.h"
> > #include "xe_pm.h"
> > @@ -281,6 +282,8 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
> > if (IS_ERR(vm))
> > return PTR_ERR(vm);
> >
> > + xe_migrate_ulls_enter(gt_to_tile(gt)->migrate);
> > +
> > down_read(&vm->lock);
> >
> > if (xe_vm_is_closed(vm)) {
> > diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> > index f66bd1de2b70..425c678f4480 100644
> > --- a/drivers/gpu/drm/xe/xe_vm.c
> > +++ b/drivers/gpu/drm/xe/xe_vm.c
> > @@ -2528,8 +2528,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> > ctx.devmem_possible = IS_DGFX(vm->xe) &&
> > IS_ENABLED(CONFIG_DRM_XE_PAGEMAP);
> >
> > - for_each_tile(tile, vm->xe, id)
> > + for_each_tile(tile, vm->xe, id) {
> > + xe_migrate_ulls_enter(tile->migrate);
> > tile_mask |= 0x1 << id;
> > + }
> >
> > if (prefetch_region == DRM_XE_CONSULT_MEM_ADVISE_PREF_LOC) {
> > dpagemap = xe_vma_resolve_pagemap(vma,
>
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
2026-09-25 18:17 ` Matthew Brost
@ 2026-09-25 18:26 ` Maarten Lankhorst
2026-09-25 19:38 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: Maarten Lankhorst @ 2026-09-25 18:26 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe, Shuicheng Lin, Himal Prasad Ghimiray
On 9/25/26 20:17, Matthew Brost wrote:
> On Fri, Sep 25, 2026 at 07:49:43PM +0200, Maarten Lankhorst wrote:
>> Hey,
>>
>> I know this has been reviewed, but can it be changed to have xe_migrate_ulls_enter + xe_migrate_ulls_leave?
>> It decreases the amount of pingponging with debugging enabled.
>>
>> I kept below patch in my tree that did just that, feel free to merge with this patch.
>>
> Ok, I see the potential races where before we get to a copy job in
> either prefetch or a pagefault before the exit period ULLS will disarm
> itself - I'd hope within the default of 5ms that would never happen in
> practice as 5ms is a *really* long time in hot paths (prefetch, SVM).
>
> Is this the issue you are seeing in debug builds, I personally have
> never noticed this but generally run with perf builds.
>
> Are you using ULLS somewhere else where 5ms can be hit more easier?
>
> Want to get a full picture before commiting to something.
It causes different behavior in a debugging kernel. PROVE_LOCKING, DEBUG_SLUB, etc may cause
a behavior change.
I believe it's better to keep it as a refcount. iirc I noticed it when investigating a test failure from C
in a previous version of your patch series.
In that testcase, the ulls fastpath was disabled before the fastpath completely ran, and a test failure happened.
I'm not sure if the latter was related, but it made me aware that it is definitely a possibility when debugging
that behavior will change compared to non-debugging case.
> Matt
>
>> Kind regards,
>> ~Maarten Lankhorst
>>
>> commit 6beeb2f7ff909bdaf9f8b1eeaa014cd15a65c78b
>> Author: Maarten Lankhorst <dev@lankhorst.se>
>> Date: Mon Jun 8 19:05:46 2026 +0200
>>
>> drm/xe: Keep ulls alive while servicing ops.
>>
>> Instead of bumping at the start, when performing a lot of work ulls may
>> run longer, so keep track using a refcount to keep the full benefit of
>> ulls.
>>
>> Signed-off-by: Maarten Lankhorst <dev@lankhorst.se>
>>
>> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
>> index d28da51d71452..f0b6641dd9815 100644
>> --- a/drivers/gpu/drm/xe/xe_migrate.c
>> +++ b/drivers/gpu/drm/xe/xe_migrate.c
>> @@ -58,6 +58,8 @@ struct xe_migrate {
>> struct xe_tile *tile;
>> /** @job_mutex: Timeline mutex for @eng. */
>> struct mutex job_mutex;
>> + /** @ulls_used: How many outstanding ulls operations there are */
>> + atomic_t ulls_used;
>> /** @pt_bo: Page-table buffer object. */
>> struct xe_bo *pt_bo;
>> /** @batch_base_ofs: VM offset of the migration batch buffer */
>> @@ -475,6 +477,8 @@ static int xe_migrate_lock_prepare_vm(struct xe_tile *tile, struct xe_migrate *m
>> *
>> * If DGFX and not a VF, enter ULLS mode bypassing GuC / HW context
>> * switches by utilizing semaphore and continuously running batches.
>> + *
>> + * Pairs with xe_migrate_ulls_leave().
>> */
>> void xe_migrate_ulls_enter(struct xe_migrate *m)
>> {
>> @@ -488,6 +492,12 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
>> if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
>> return;
>>
>> + /* Bump */
>> + if (atomic_inc_return(&m->ulls_used) > 1)
>> + return;
>> +
>> + cancel_delayed_work_sync(&m->ulls.exit_work);
>> +
>> job_alloc:
>> if (alloc) {
>> /*
>> @@ -534,10 +544,30 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
>> }
>> if (job)
>> xe_sched_job_put(job);
>> + mutex_unlock(&m->job_mutex);
>> +}
>> +
>> +/**
>> + * xe_migrate_ulls_leave() - Leave ULLS mode
>> + * @m: The migration context.
>> + *
>> + * Leaves the critical part of the migration context, it may continue to be enabled.
>> + * Pairs with xe_migrate_ulls_enter().
>> + */
>> +void xe_migrate_ulls_leave(struct xe_migrate *m)
>> +{
>> + struct xe_device *xe = tile_to_xe(m->tile);
>> +
>> + if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
>> + return;
>> +
>> + if (atomic_dec_return(&m->ulls_used))
>> + return;
>> +
>> + guard(mutex)(&m->job_mutex);
>> if (m->ulls.enabled)
>> mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
>> ULLS_EXIT_JIFFIES);
>> - mutex_unlock(&m->job_mutex);
>> }
>>
>> static void xe_migrate_ulls_exit(struct work_struct *work)
>> @@ -599,8 +629,8 @@ static void xe_migrate_ulls_exit(struct work_struct *work)
>> ULLS_EXIT_JIFFIES);
>> }
>>
>> - drm_dev_exit(idx);
>> mutex_unlock(&m->job_mutex);
>> + drm_dev_exit(idx);
>> }
>>
>> /**
>> diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
>> index 2e8be10fdb71f..c909e1afb476e 100644
>> --- a/drivers/gpu/drm/xe/xe_migrate.h
>> +++ b/drivers/gpu/drm/xe/xe_migrate.h
>> @@ -98,5 +98,6 @@ int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
>> #endif
>>
>> void xe_migrate_ulls_enter(struct xe_migrate *m);
>> +void xe_migrate_ulls_leave(struct xe_migrate *m);
>>
>> #endif
>> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
>> index 782ed50a2a43e..51b2cb8fdcf1c 100644
>> --- a/drivers/gpu/drm/xe/xe_pagefault.c
>> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
>> @@ -315,6 +315,7 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
>>
>> unlock_vm:
>> up_read(&vm->lock);
>> + xe_migrate_ulls_leave(gt_to_tile(gt)->migrate);
>> xe_vm_put(vm);
>>
>> return err;
>> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
>> index a1eca4871175b..fc9f587a35905 100644
>> --- a/drivers/gpu/drm/xe/xe_vm.c
>> +++ b/drivers/gpu/drm/xe/xe_vm.c
>> @@ -2557,7 +2557,7 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
>> if (addr)
>> goto alloc_next_range;
>> else
>> - goto print_op_label;
>> + goto ulls_leave;
>> }
>>
>> if (IS_ERR(svm_range)) {
>> @@ -2598,8 +2598,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
>> if (need_put)
>> xe_svm_range_put(svm_range);
>>
>> +ulls_leave:
>> + for_each_tile(tile, vm->xe, id)
>> + xe_migrate_ulls_leave(tile->migrate);
>> }
>> -print_op_label:
>> print_op(vm->xe, __op);
>> }
>>
>>
>>
>>
>> On 9/25/26 06:53, Matthew Brost wrote:
>>> Call xe_migrate_ulls_enter upon page fault or SVM prefetch in an
>>> effort speed up these critical paths.
>>>
>>> Signed-off-by: Matthew Brost <matthew.brost@intel.com>
>>> Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
>>> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
>>>
>>> ---
>>> v7:
>>> - s/migration/migrate (Shuicheng)
>>> ---
>>> drivers/gpu/drm/xe/xe_pagefault.c | 3 +++
>>> drivers/gpu/drm/xe/xe_vm.c | 4 +++-
>>> 2 files changed, 6 insertions(+), 1 deletion(-)
>>>
>>> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
>>> index f9e5bd41cfd1..f219e9e3b73e 100644
>>> --- a/drivers/gpu/drm/xe/xe_pagefault.c
>>> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
>>> @@ -15,6 +15,7 @@
>>> #include "xe_gt_stats.h"
>>> #include "xe_hw_engine.h"
>>> #include "xe_log.h"
>>> +#include "xe_migrate.h"
>>> #include "xe_pagefault.h"
>>> #include "xe_pagefault_types.h"
>>> #include "xe_pm.h"
>>> @@ -281,6 +282,8 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
>>> if (IS_ERR(vm))
>>> return PTR_ERR(vm);
>>>
>>> + xe_migrate_ulls_enter(gt_to_tile(gt)->migrate);
>>> +
>>> down_read(&vm->lock);
>>>
>>> if (xe_vm_is_closed(vm)) {
>>> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
>>> index f66bd1de2b70..425c678f4480 100644
>>> --- a/drivers/gpu/drm/xe/xe_vm.c
>>> +++ b/drivers/gpu/drm/xe/xe_vm.c
>>> @@ -2528,8 +2528,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
>>> ctx.devmem_possible = IS_DGFX(vm->xe) &&
>>> IS_ENABLED(CONFIG_DRM_XE_PAGEMAP);
>>>
>>> - for_each_tile(tile, vm->xe, id)
>>> + for_each_tile(tile, vm->xe, id) {
>>> + xe_migrate_ulls_enter(tile->migrate);
>>> tile_mask |= 0x1 << id;
>>> + }
>>>
>>> if (prefetch_region == DRM_XE_CONSULT_MEM_ADVISE_PREF_LOC) {
>>> dpagemap = xe_vma_resolve_pagemap(vma,
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
2026-09-25 18:26 ` Maarten Lankhorst
@ 2026-09-25 19:38 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 19:38 UTC (permalink / raw)
To: Maarten Lankhorst; +Cc: intel-xe, Shuicheng Lin, Himal Prasad Ghimiray
On Fri, Sep 25, 2026 at 08:26:27PM +0200, Maarten Lankhorst wrote:
>
>
> On 9/25/26 20:17, Matthew Brost wrote:
> > On Fri, Sep 25, 2026 at 07:49:43PM +0200, Maarten Lankhorst wrote:
> >> Hey,
> >>
> >> I know this has been reviewed, but can it be changed to have xe_migrate_ulls_enter + xe_migrate_ulls_leave?
> >> It decreases the amount of pingponging with debugging enabled.
> >>
> >> I kept below patch in my tree that did just that, feel free to merge with this patch.
> >>
> > Ok, I see the potential races where before we get to a copy job in
> > either prefetch or a pagefault before the exit period ULLS will disarm
> > itself - I'd hope within the default of 5ms that would never happen in
> > practice as 5ms is a *really* long time in hot paths (prefetch, SVM).
> >
> > Is this the issue you are seeing in debug builds, I personally have
> > never noticed this but generally run with perf builds.
> >
> > Are you using ULLS somewhere else where 5ms can be hit more easier?
> >
> > Want to get a full picture before commiting to something.
> It causes different behavior in a debugging kernel. PROVE_LOCKING, DEBUG_SLUB, etc may cause
> a behavior change.
>
> I believe it's better to keep it as a refcount. iirc I noticed it when investigating a test failure from C
> in a previous version of your patch series.
>
> In that testcase, the ulls fastpath was disabled before the fastpath completely ran, and a test failure happened.
> I'm not sure if the latter was related, but it made me aware that it is definitely a possibility when debugging
> that behavior will change compared to non-debugging case.
Well ping-ponging should work, I probably had some bugs in earlier revs
but let's me force ping-pong now and ensure that works.
Can we do this is a follow up - not opposed to the idea and we will
probably have use cases in future where we want a longterm pin on ULLS.
Matt
>
> > Matt
> >
> >> Kind regards,
> >> ~Maarten Lankhorst
> >>
> >> commit 6beeb2f7ff909bdaf9f8b1eeaa014cd15a65c78b
> >> Author: Maarten Lankhorst <dev@lankhorst.se>
> >> Date: Mon Jun 8 19:05:46 2026 +0200
> >>
> >> drm/xe: Keep ulls alive while servicing ops.
> >>
> >> Instead of bumping at the start, when performing a lot of work ulls may
> >> run longer, so keep track using a refcount to keep the full benefit of
> >> ulls.
> >>
> >> Signed-off-by: Maarten Lankhorst <dev@lankhorst.se>
> >>
> >> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> >> index d28da51d71452..f0b6641dd9815 100644
> >> --- a/drivers/gpu/drm/xe/xe_migrate.c
> >> +++ b/drivers/gpu/drm/xe/xe_migrate.c
> >> @@ -58,6 +58,8 @@ struct xe_migrate {
> >> struct xe_tile *tile;
> >> /** @job_mutex: Timeline mutex for @eng. */
> >> struct mutex job_mutex;
> >> + /** @ulls_used: How many outstanding ulls operations there are */
> >> + atomic_t ulls_used;
> >> /** @pt_bo: Page-table buffer object. */
> >> struct xe_bo *pt_bo;
> >> /** @batch_base_ofs: VM offset of the migration batch buffer */
> >> @@ -475,6 +477,8 @@ static int xe_migrate_lock_prepare_vm(struct xe_tile *tile, struct xe_migrate *m
> >> *
> >> * If DGFX and not a VF, enter ULLS mode bypassing GuC / HW context
> >> * switches by utilizing semaphore and continuously running batches.
> >> + *
> >> + * Pairs with xe_migrate_ulls_leave().
> >> */
> >> void xe_migrate_ulls_enter(struct xe_migrate *m)
> >> {
> >> @@ -488,6 +492,12 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
> >> if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
> >> return;
> >>
> >> + /* Bump */
> >> + if (atomic_inc_return(&m->ulls_used) > 1)
> >> + return;
> >> +
> >> + cancel_delayed_work_sync(&m->ulls.exit_work);
> >> +
> >> job_alloc:
> >> if (alloc) {
> >> /*
> >> @@ -534,10 +544,30 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
> >> }
> >> if (job)
> >> xe_sched_job_put(job);
> >> + mutex_unlock(&m->job_mutex);
> >> +}
> >> +
> >> +/**
> >> + * xe_migrate_ulls_leave() - Leave ULLS mode
> >> + * @m: The migration context.
> >> + *
> >> + * Leaves the critical part of the migration context, it may continue to be enabled.
> >> + * Pairs with xe_migrate_ulls_enter().
> >> + */
> >> +void xe_migrate_ulls_leave(struct xe_migrate *m)
> >> +{
> >> + struct xe_device *xe = tile_to_xe(m->tile);
> >> +
> >> + if (!IS_DGFX(xe) || IS_SRIOV_VF(xe) || !xe->info.ulls_enable)
> >> + return;
> >> +
> >> + if (atomic_dec_return(&m->ulls_used))
> >> + return;
> >> +
> >> + guard(mutex)(&m->job_mutex);
> >> if (m->ulls.enabled)
> >> mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
> >> ULLS_EXIT_JIFFIES);
> >> - mutex_unlock(&m->job_mutex);
> >> }
> >>
> >> static void xe_migrate_ulls_exit(struct work_struct *work)
> >> @@ -599,8 +629,8 @@ static void xe_migrate_ulls_exit(struct work_struct *work)
> >> ULLS_EXIT_JIFFIES);
> >> }
> >>
> >> - drm_dev_exit(idx);
> >> mutex_unlock(&m->job_mutex);
> >> + drm_dev_exit(idx);
> >> }
> >>
> >> /**
> >> diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
> >> index 2e8be10fdb71f..c909e1afb476e 100644
> >> --- a/drivers/gpu/drm/xe/xe_migrate.h
> >> +++ b/drivers/gpu/drm/xe/xe_migrate.h
> >> @@ -98,5 +98,6 @@ int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
> >> #endif
> >>
> >> void xe_migrate_ulls_enter(struct xe_migrate *m);
> >> +void xe_migrate_ulls_leave(struct xe_migrate *m);
> >>
> >> #endif
> >> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
> >> index 782ed50a2a43e..51b2cb8fdcf1c 100644
> >> --- a/drivers/gpu/drm/xe/xe_pagefault.c
> >> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
> >> @@ -315,6 +315,7 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
> >>
> >> unlock_vm:
> >> up_read(&vm->lock);
> >> + xe_migrate_ulls_leave(gt_to_tile(gt)->migrate);
> >> xe_vm_put(vm);
> >>
> >> return err;
> >> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> >> index a1eca4871175b..fc9f587a35905 100644
> >> --- a/drivers/gpu/drm/xe/xe_vm.c
> >> +++ b/drivers/gpu/drm/xe/xe_vm.c
> >> @@ -2557,7 +2557,7 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> >> if (addr)
> >> goto alloc_next_range;
> >> else
> >> - goto print_op_label;
> >> + goto ulls_leave;
> >> }
> >>
> >> if (IS_ERR(svm_range)) {
> >> @@ -2598,8 +2598,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> >> if (need_put)
> >> xe_svm_range_put(svm_range);
> >>
> >> +ulls_leave:
> >> + for_each_tile(tile, vm->xe, id)
> >> + xe_migrate_ulls_leave(tile->migrate);
> >> }
> >> -print_op_label:
> >> print_op(vm->xe, __op);
> >> }
> >>
> >>
> >>
> >>
> >> On 9/25/26 06:53, Matthew Brost wrote:
> >>> Call xe_migrate_ulls_enter upon page fault or SVM prefetch in an
> >>> effort speed up these critical paths.
> >>>
> >>> Signed-off-by: Matthew Brost <matthew.brost@intel.com>
> >>> Reviewed-by: Shuicheng Lin <shuicheng.lin@intel.com>
> >>> Reviewed-by: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> >>>
> >>> ---
> >>> v7:
> >>> - s/migration/migrate (Shuicheng)
> >>> ---
> >>> drivers/gpu/drm/xe/xe_pagefault.c | 3 +++
> >>> drivers/gpu/drm/xe/xe_vm.c | 4 +++-
> >>> 2 files changed, 6 insertions(+), 1 deletion(-)
> >>>
> >>> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
> >>> index f9e5bd41cfd1..f219e9e3b73e 100644
> >>> --- a/drivers/gpu/drm/xe/xe_pagefault.c
> >>> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
> >>> @@ -15,6 +15,7 @@
> >>> #include "xe_gt_stats.h"
> >>> #include "xe_hw_engine.h"
> >>> #include "xe_log.h"
> >>> +#include "xe_migrate.h"
> >>> #include "xe_pagefault.h"
> >>> #include "xe_pagefault_types.h"
> >>> #include "xe_pm.h"
> >>> @@ -281,6 +282,8 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
> >>> if (IS_ERR(vm))
> >>> return PTR_ERR(vm);
> >>>
> >>> + xe_migrate_ulls_enter(gt_to_tile(gt)->migrate);
> >>> +
> >>> down_read(&vm->lock);
> >>>
> >>> if (xe_vm_is_closed(vm)) {
> >>> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> >>> index f66bd1de2b70..425c678f4480 100644
> >>> --- a/drivers/gpu/drm/xe/xe_vm.c
> >>> +++ b/drivers/gpu/drm/xe/xe_vm.c
> >>> @@ -2528,8 +2528,10 @@ vm_bind_ioctl_ops_create(struct xe_vm *vm, struct xe_vma_ops *vops,
> >>> ctx.devmem_possible = IS_DGFX(vm->xe) &&
> >>> IS_ENABLED(CONFIG_DRM_XE_PAGEMAP);
> >>>
> >>> - for_each_tile(tile, vm->xe, id)
> >>> + for_each_tile(tile, vm->xe, id) {
> >>> + xe_migrate_ulls_enter(tile->migrate);
> >>> tile_mask |= 0x1 << id;
> >>> + }
> >>>
> >>> if (prefetch_region == DRM_XE_CONSULT_MEM_ADVISE_PREF_LOC) {
> >>> dpagemap = xe_vma_resolve_pagemap(vma,
>
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (21 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 6:51 ` sashiko-bot
2026-09-25 4:53 ` [PATCH v7 24/24] drm/xe: Document ULLS for migration jobs Matthew Brost
` (5 subsequent siblings)
28 siblings, 1 reply; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe
Add a configfs attribute to control the migrate ULLS period.
The migrate ULLS period specifies how long, in milliseconds, migrate
ULLS remains active before exiting. Migrate ULLS is currently used by
page-fault handling and SVM prefetch operations to process migration
work through a continuously running batch.
The attribute defaults to 5 ms. Setting the value to 0 disables ULLS.
This knob is primarily intended for performance tuning and evaluation,
allowing experimentation with the trade-off between reducing ULLS
startup/teardown overhead and limiting the time spent in a continuously
running migration workload.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
---
drivers/gpu/drm/xe/xe_configfs.c | 67 ++++++++++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_configfs.h | 2 +
drivers/gpu/drm/xe/xe_migrate.c | 19 ++++++---
3 files changed, 82 insertions(+), 6 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_configfs.c b/drivers/gpu/drm/xe/xe_configfs.c
index f5c828cf7e8f..f23761f2eb11 100644
--- a/drivers/gpu/drm/xe/xe_configfs.c
+++ b/drivers/gpu/drm/xe/xe_configfs.c
@@ -61,6 +61,7 @@
* ├── survivability_mode
* ├── gt_types_allowed
* ├── engines_allowed
+ * ├── migrate_ulls_period_ms
* ├── enable_psmi
* └── disable_vram_page_offline
*
@@ -262,6 +263,20 @@
*
* This attribute can only be set before binding to the device.
*
+ * Migrate ULLS Period (ms)
+ * ------------------------
+ *
+ * Migrate ULLS period, in milliseconds. This is the delay between entering
+ * migrate ULLS (a continuously running batch) and exiting it. Migrate ULLS is
+ * currently entered during page faults and SVM prefetch operations. Default 5,
+ * zero indicates ULLS is disabled.
+ *
+ * How to disable migration ULLS:
+ *
+ * # echo 0 > /sys/kernel/config/xe/0000:03:00.0/migrate_ulls_period_ms
+ *
+ * This attribute can only be set before binding to the device.
+ *
* Remove devices
* ==============
*
@@ -283,6 +298,7 @@ struct xe_config_group_device {
struct xe_config_device {
u64 gt_types_allowed;
u64 engines_allowed;
+ u32 migrate_ulls_period_ms;
struct wa_bb ctx_restore_post_bb[XE_ENGINE_CLASS_MAX];
struct wa_bb ctx_restore_mid_bb[XE_ENGINE_CLASS_MAX];
bool survivability_mode;
@@ -306,6 +322,7 @@ struct xe_config_group_device {
static const struct xe_config_device device_defaults = {
.gt_types_allowed = U64_MAX,
.engines_allowed = U64_MAX,
+ .migrate_ulls_period_ms = 5,
.survivability_mode = false,
.enable_psmi = false,
.enable_multi_queue = true,
@@ -631,6 +648,33 @@ static ssize_t enable_multi_queue_store(struct config_item *item, const char *pa
return len;
}
+static ssize_t migrate_ulls_period_ms_show(struct config_item *item, char *page)
+{
+ struct xe_config_device *dev = to_xe_config_device(item);
+
+ return sprintf(page, "%d\n", dev->migrate_ulls_period_ms);
+}
+
+static ssize_t migrate_ulls_period_ms_store(struct config_item *item,
+ const char *page, size_t len)
+{
+ struct xe_config_group_device *dev = to_xe_config_group_device(item);
+ u32 val;
+ int ret;
+
+ ret = kstrtou32(page, 0, &val);
+ if (ret)
+ return ret;
+
+ guard(mutex)(&dev->lock);
+ if (is_bound(dev))
+ return -EBUSY;
+
+ dev->config.migrate_ulls_period_ms = val;
+
+ return len;
+}
+
static ssize_t disable_vram_page_offline_show(struct config_item *item, char *page)
{
struct xe_config_device *dev = to_xe_config_device(item);
@@ -896,6 +940,7 @@ static ssize_t ctx_restore_post_bb_store(struct config_item *item,
CONFIGFS_ATTR(, ctx_restore_mid_bb);
CONFIGFS_ATTR(, ctx_restore_post_bb);
CONFIGFS_ATTR(, enable_multi_queue);
+CONFIGFS_ATTR(, migrate_ulls_period_ms);
CONFIGFS_ATTR(, enable_psmi);
CONFIGFS_ATTR(, disable_vram_page_offline);
CONFIGFS_ATTR(, engines_allowed);
@@ -906,6 +951,7 @@ static struct configfs_attribute *xe_config_device_attrs[] = {
&attr_ctx_restore_mid_bb,
&attr_ctx_restore_post_bb,
&attr_enable_multi_queue,
+ &attr_migrate_ulls_period_ms,
&attr_enable_psmi,
&attr_disable_vram_page_offline,
&attr_engines_allowed,
@@ -1189,6 +1235,7 @@ static void dump_custom_dev_config(struct pci_dev *pdev,
PRI_CUSTOM_ATTR("%llx", gt_types_allowed);
PRI_CUSTOM_ATTR("%llx", engines_allowed);
+ PRI_CUSTOM_ATTR("%u", migrate_ulls_period_ms);
PRI_CUSTOM_ATTR("%d", enable_multi_queue);
PRI_CUSTOM_ATTR("%d", enable_psmi);
PRI_CUSTOM_ATTR("%d", disable_vram_page_offline);
@@ -1340,6 +1387,26 @@ bool xe_configfs_get_enable_multi_queue(struct pci_dev *pdev)
return ret;
}
+/**
+ * xe_configfs_get_migrate_ulls_period_ms - get configfs migrate_ulls_period_ms setting
+ * @pdev: pci device
+ *
+ * Return: Migrate ULLS period in milliseconds (zero is disabled).
+ */
+u32 xe_configfs_get_migrate_ulls_period_ms(struct pci_dev *pdev)
+{
+ struct xe_config_group_device *dev = find_xe_config_group_device(pdev);
+ u32 ret;
+
+ if (!dev)
+ return device_defaults.migrate_ulls_period_ms;
+
+ ret = dev->config.migrate_ulls_period_ms;
+ config_group_put(&dev->group);
+
+ return ret;
+}
+
/**
* xe_configfs_get_disable_vram_page_offline - get configfs disable_vram_page_offline setting
* @pdev: pci device
diff --git a/drivers/gpu/drm/xe/xe_configfs.h b/drivers/gpu/drm/xe/xe_configfs.h
index 42cd1a491d01..10a00d6bbf8f 100644
--- a/drivers/gpu/drm/xe/xe_configfs.h
+++ b/drivers/gpu/drm/xe/xe_configfs.h
@@ -24,6 +24,7 @@ bool xe_configfs_media_gt_allowed(struct pci_dev *pdev);
u64 xe_configfs_get_engines_allowed(struct pci_dev *pdev);
bool xe_configfs_get_psmi_enabled(struct pci_dev *pdev);
bool xe_configfs_get_enable_multi_queue(struct pci_dev *pdev);
+u32 xe_configfs_get_migrate_ulls_period_ms(struct pci_dev *pdev);
bool xe_configfs_get_disable_vram_page_offline(struct pci_dev *pdev);
u32 xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev,
enum xe_engine_class class,
@@ -45,6 +46,7 @@ static inline bool xe_configfs_media_gt_allowed(struct pci_dev *pdev) { return t
static inline u64 xe_configfs_get_engines_allowed(struct pci_dev *pdev) { return U64_MAX; }
static inline bool xe_configfs_get_psmi_enabled(struct pci_dev *pdev) { return false; }
static inline bool xe_configfs_get_enable_multi_queue(struct pci_dev *pdev) { return true; }
+static inline u32 xe_configfs_get_migrate_ulls_period_ms(struct pci_dev *pdev) { return 5; }
static inline bool xe_configfs_get_disable_vram_page_offline(struct pci_dev *pdev) { return false; }
static inline u32 xe_configfs_get_ctx_restore_mid_bb(struct pci_dev *pdev,
enum xe_engine_class class,
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index d7d13a25cdb9..c643177ff2d0 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -23,6 +23,7 @@
#include "xe_assert.h"
#include "xe_bb.h"
#include "xe_bo.h"
+#include "xe_configfs.h"
#include "xe_exec_queue.h"
#include "xe_ggtt.h"
#include "xe_gt.h"
@@ -81,14 +82,17 @@ struct xe_migrate {
u64 min_chunk_size;
/** @ulls: ULLS support */
struct {
+ /** @ulls.exit_ms: ULLS exit period milliseconds */
+ u32 exit_ms;
/** @ulls.enabled: ULLS is enabled, protected by job_mutex */
bool enabled;
-#define ULLS_EXIT_JIFFIES msecs_to_jiffies(5)
/** @ulls.exit_work: ULLS exit worker */
struct delayed_work exit_work;
} ulls;
};
+#define ULLS_EXIT_JIFFIES(_m) msecs_to_jiffies((_m)->ulls.exit_ms)
+
#define MAX_PREEMPTDISABLE_TRANSFER SZ_8M /* Around 1ms. */
#define MAX_CCS_LIMITED_TRANSFER SZ_4M /* XE_PAGE_SIZE * (FIELD_MAX(XE2_CCS_SIZE_MASK) + 1) */
#define NUM_PT_SLOTS 48
@@ -512,7 +516,7 @@ static struct dma_fence *xe_migrate_job_push(struct xe_migrate *m,
if (xe_migrate_ulls_enabled(m)) {
ulls = ULLS_ACTIVE;
mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
- ULLS_EXIT_JIFFIES);
+ ULLS_EXIT_JIFFIES(m));
}
return __xe_migrate_job_push(m, job, ulls);
@@ -534,7 +538,7 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
xe_assert(xe, xe->info.has_usm);
- if (!IS_DGFX(xe))
+ if (!IS_DGFX(xe) || !m->ulls.exit_ms)
return;
job_alloc:
@@ -572,7 +576,7 @@ void xe_migrate_ulls_enter(struct xe_migrate *m)
xe_sched_job_put(job);
if (xe_migrate_ulls_enabled(m))
mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
- ULLS_EXIT_JIFFIES);
+ ULLS_EXIT_JIFFIES(m));
mutex_unlock(&m->job_mutex);
}
@@ -599,7 +603,7 @@ static void xe_migrate_ulls_exit(struct work_struct *work)
if (WARN_ON_ONCE(IS_ERR(job))) {
drm_dev_exit(idx);
mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
- ULLS_EXIT_JIFFIES);
+ ULLS_EXIT_JIFFIES(m));
return; /* Not fatal */
}
@@ -616,7 +620,7 @@ static void xe_migrate_ulls_exit(struct work_struct *work)
} else {
xe_sched_job_put(job);
mod_delayed_work(system_percpu_wq, &m->ulls.exit_work,
- ULLS_EXIT_JIFFIES);
+ ULLS_EXIT_JIFFIES(m));
}
}
@@ -700,6 +704,9 @@ int xe_migrate_init(struct xe_migrate *m)
(unsigned long long)m->min_chunk_size);
}
+ m->ulls.exit_ms =
+ xe_configfs_get_migrate_ulls_period_ms(to_pci_dev(xe->drm.dev));
+
return err;
err_out:
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* Re: [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute
2026-09-25 4:53 ` [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute Matthew Brost
@ 2026-09-25 6:51 ` sashiko-bot
2026-09-25 7:08 ` Matthew Brost
0 siblings, 1 reply; 54+ messages in thread
From: sashiko-bot @ 2026-09-25 6:51 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] xe: format string mismatch for u32 configfs attribute
--
commit 11e688bb0c4319099fe2944add90b93a411505ab
Author: Matthew Brost <matthew.brost@intel.com>
drm/xe: add migrate ULLS period configfs attribute
This commit adds a configfs attribute to control the migrate ULLS period.
The period specifies how long in milliseconds migrate ULLS remains active
before exiting, primarily intended for performance tuning.
> diff --git a/drivers/gpu/drm/xe/xe_configfs.c b/drivers/gpu/drm/xe/xe_configfs.c
> index f1555a9d31170..3f09dd17f80e0 100644
> --- a/drivers/gpu/drm/xe/xe_configfs.c
> +++ b/drivers/gpu/drm/xe/xe_configfs.c
[ ... ]
> @@ -631,6 +648,33 @@ static ssize_t enable_multi_queue_store(struct config_item *item, const char *pa
> return len;
> }
>
> +static ssize_t migrate_ulls_period_ms_show(struct config_item *item, char *page)
> +{
> + struct xe_config_device *dev = to_xe_config_device(item);
> +
> + return sprintf(page, "%d\n", dev->migrate_ulls_period_ms);
[Severity: Low]
Since dev->migrate_ulls_period_ms is an unsigned 32-bit integer (u32),
could writing a value greater than INT_MAX via configfs cause it to be read
back as a negative number due to the "%d" format string?
Would "%u" be more appropriate here to preserve the read/write round-trip
property for userspace?
> +}
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=23
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute
2026-09-25 6:51 ` sashiko-bot
@ 2026-09-25 7:08 ` Matthew Brost
0 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 7:08 UTC (permalink / raw)
To: sashiko-reviews; +Cc: intel-xe
On Fri, Sep 25, 2026 at 06:51:54AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Low] xe: format string mismatch for u32 configfs attribute
> --
>
> commit 11e688bb0c4319099fe2944add90b93a411505ab
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: add migrate ULLS period configfs attribute
>
> This commit adds a configfs attribute to control the migrate ULLS period.
> The period specifies how long in milliseconds migrate ULLS remains active
> before exiting, primarily intended for performance tuning.
>
> > diff --git a/drivers/gpu/drm/xe/xe_configfs.c b/drivers/gpu/drm/xe/xe_configfs.c
> > index f1555a9d31170..3f09dd17f80e0 100644
> > --- a/drivers/gpu/drm/xe/xe_configfs.c
> > +++ b/drivers/gpu/drm/xe/xe_configfs.c
> [ ... ]
> > @@ -631,6 +648,33 @@ static ssize_t enable_multi_queue_store(struct config_item *item, const char *pa
> > return len;
> > }
> >
> > +static ssize_t migrate_ulls_period_ms_show(struct config_item *item, char *page)
> > +{
> > + struct xe_config_device *dev = to_xe_config_device(item);
> > +
> > + return sprintf(page, "%d\n", dev->migrate_ulls_period_ms);
>
> [Severity: Low]
> Since dev->migrate_ulls_period_ms is an unsigned 32-bit integer (u32),
> could writing a value greater than INT_MAX via configfs cause it to be read
> back as a negative number due to the "%d" format string?
>
> Would "%u" be more appropriate here to preserve the read/write round-trip
> property for userspace?
Yes, %u is correct. Will fix.
Matt
>
> > +}
> [ ... ]
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=23
^ permalink raw reply [flat|nested] 54+ messages in thread
* [PATCH v7 24/24] drm/xe: Document ULLS for migration jobs
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (22 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute Matthew Brost
@ 2026-09-25 4:53 ` Matthew Brost
2026-09-25 5:02 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev9) Patchwork
` (4 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Matthew Brost @ 2026-09-25 4:53 UTC (permalink / raw)
To: intel-xe
Add a kernel-doc DOC section at the top of xe_migrate.c describing the
Ultra Low Latency Submission (ULLS) scheme used for migration jobs.
Cover the motivation (removing the H2G / GuC / context switch latency
from the page fault and SVM prefetch critical paths), the platform
requirements, the LRC PPHWSP semaphore layout and its relationship to
the migration queue job count, the fixed ULLS job size and why it is
needed, the ring preamble / postamble emitted by the ring ops including
the in-ring tail update, the semaphore-only submission fast path in the
GuC backend, and the enter / delayed exit flow along with the ULLS job
flags.
Hook the new section into Documentation/gpu/xe/xe_migrate.rst.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Assisted-by: Github-Copilot:Claude-opus-5
---
v7:
- configfs control (Thomas)
---
Documentation/gpu/xe/xe_migrate.rst | 3 +
drivers/gpu/drm/xe/xe_migrate.c | 149 ++++++++++++++++++++++++++++
2 files changed, 152 insertions(+)
diff --git a/Documentation/gpu/xe/xe_migrate.rst b/Documentation/gpu/xe/xe_migrate.rst
index f92faec0ac94..d297ee53a582 100644
--- a/Documentation/gpu/xe/xe_migrate.rst
+++ b/Documentation/gpu/xe/xe_migrate.rst
@@ -6,3 +6,6 @@ Migrate Layer
.. kernel-doc:: drivers/gpu/drm/xe/xe_migrate_doc.h
:doc: Migrate Layer
+
+.. kernel-doc:: drivers/gpu/drm/xe/xe_migrate.c
+ :doc: ULLS (Ultra Low Latency Submission) for migration jobs
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index c643177ff2d0..0dfc54ba3b8f 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -48,6 +48,155 @@
#include "xe_vm.h"
#include "xe_vram.h"
+/**
+ * DOC: ULLS (Ultra Low Latency Submission) for migration jobs
+ *
+ * Migration jobs issued on behalf of GPU page faults and SVM prefetches sit
+ * directly in the critical path of a stalled GPU workload. The dominant cost
+ * of such a job is not the copy or clear itself but the submission latency:
+ * the H2G round trip to GuC, the GuC scheduling decision, and the hardware
+ * context switch required to place the migration LRC on an engine.
+ *
+ * ULLS removes that cost by keeping the migration context resident and
+ * *running* on the hardware engine across jobs. Instead of the ring going
+ * empty and the context being switched out between jobs, the tail of every
+ * ULLS job parks the engine on a semaphore wait for the *next* job's
+ * semaphore, and then advances the ring tail itself. Submitting the next job
+ * therefore costs the CPU a single write to signal that semaphore - no H2G,
+ * no GuC round trip, no context switch, no MMIO.
+ *
+ * Requirements
+ * ------------
+ *
+ * ULLS is only used on dGFX platforms with USM support, where a hardware
+ * engine is reserved exclusively for migration jobs. Because the engine
+ * spins on a semaphore while ULLS is active, it cannot be shared with
+ * user submissions.
+ *
+ * ULLS can be disabled by setting the ``migrate_ulls_period_ms`` configfs
+ * attribute to 0. Otherwise, the same attribute controls how long ULLS
+ * remains active before exiting, in milliseconds.
+ *
+ * Fixed size jobs
+ * ---------------
+ *
+ * A job updates the ring tail to cover its successor, but it is emitted long
+ * before that successor exists, so it can not know how much ring the
+ * successor will occupy. Every ULLS job is therefore padded out to exactly
+ * ULLS_JOB_SIZE_BYTES, which lets the next tail be computed arithmetically
+ * from where the current job started.
+ *
+ * This is why the shorter jobs still have to reach the same size: the "last"
+ * job skips the batch buffers and the postamble, and pads the difference with
+ * MI_NOOP. The "first" job is not covered by any predecessor's tail update
+ * and so is unconstrained, but is padded anyway to keep the arithmetic
+ * uniform.
+ *
+ * Leaving ULLS mode always goes through a "last" job, which emits no tail
+ * update, so an ordinary variable length migration job never follows a
+ * prediction.
+ *
+ * Semaphores
+ * ----------
+ *
+ * The semaphores live in the driver-defined portion of the migration LRC's
+ * PPHWSP (see LRC_ULLS_PPHWSP_OFFSET, mutually exclusive with the parallel
+ * submission area). There are LRC_MIGRATION_ULLS_SEMAPHORE_COUNT of them and
+ * a job's semaphore is selected by ``seqno % COUNT``, so the semaphore ring
+ * wraps with the job seqnos. To guarantee a job can never overwrite the
+ * semaphore of a job still in flight, the GuC backend caps the migration
+ * queue's scheduler job count at LRC_MIGRATION_ULLS_SEMAPHORE_COUNT - 1.
+ *
+ * Ring layout of a ULLS job
+ * -------------------------
+ *
+ * Emitted by emit_migration_job_gen12() in xe_ring_ops.c::
+ *
+ * preamble: clear semaphore[seqno] (reuse for a later wrap)
+ * <copy timestamp, start seqno store>
+ * <batch buffer start(s)> (skipped on first/last job)
+ * <seqno write + user interrupt>
+ * postamble: SDI saved ring tail = end of next job
+ * LRI RING_TAIL = end of next job
+ * wait on semaphore[seqno + 1]
+ * (skipped on the last job)
+ * pad: MI_NOOP up to ULLS_JOB_SIZE_DW
+ *
+ * The preamble clears the current job's semaphore so it can be reused once
+ * the seqno space wraps. The postamble is what keeps the engine busy: it
+ * advances the ring tail over the next job and then blocks on that job's
+ * semaphore, which is only signaled when the job is actually submitted. It
+ * advances the saved tail as well as the tail register, keeping the two in
+ * step without any help from the CPU, so a context save and restore can not
+ * rewind the tail behind work which has already been published.
+ *
+ * The tail register write must be non-posted, i.e. it must not carry
+ * MI_LRI_FORCE_POSTED. Posted, the new tail is free to land after the command
+ * streamer has already drained the rest of the job, at which point the command
+ * streamer sees head == the old tail and parks as though the ring were empty.
+ * A parked context can be switched off the hardware, and the fast path below
+ * has no H2G with which to ask GuC to bring it back.
+ *
+ * The tail is published ahead of the semaphore wait rather than after it so
+ * that the non-posted write drains while the engine is parked anyway, keeping
+ * a register round trip off the path between the semaphore being signaled and
+ * the next job running.
+ *
+ * Submission fast path
+ * --------------------
+ *
+ * In submit_exec_queue() (xe_guc_submit.c), a ULLS job that is not the first
+ * one reduces to::
+ *
+ * xe_lrc_set_ulls_semaphore(lrc, seqno); release previous job
+ *
+ * The XE_GUC_ACTION_SCHED_CONTEXT H2G is suppressed, and so is the write of
+ * the saved ring tail: the previous job's postamble has already published
+ * this job's tail both in the tail register and in the context image, so the
+ * semaphore signal is all that is left. The previous job's semaphore wait is
+ * satisfied and the engine walks straight into this job.
+ *
+ * This does assume the context stays resident for as long as ULLS mode is
+ * active. Nothing else is scheduled on the reserved engine, so the only ways
+ * off the hardware are the "last" job below, or a reset - and a migration job
+ * failing already wedges the device.
+ *
+ * Enter / exit
+ * ------------
+ *
+ * xe_migrate_ulls_enter() is called from the page fault handler and from the
+ * SVM prefetch path, i.e. exactly where low latency migration matters. It
+ * takes a PM runtime reference (the device must not suspend while the engine
+ * spins), then submits a "first" ULLS job. That first job carries no batch
+ * buffer; it exists only to get the context onto the hardware through the
+ * normal GuC path and to leave the engine waiting on the next semaphore,
+ * pipelining the GuC/HW context switch out of the critical path.
+ *
+ * No forcewake reference is required. Nothing in the fast path touches MMIO,
+ * and the engine keeps itself awake for as long as it is executing the ring.
+ * Not needing host MMIO access is also what lets ULLS run on SRIOV VFs.
+ *
+ * Keeping an engine spinning costs power, so ULLS is not left enabled
+ * indefinitely. Every enter and every ULLS job submission re-arms
+ * @xe_migrate.ulls.exit_work with a ULLS_EXIT_JIFFIES delay. When it fires
+ * with the queue idle, it submits a "last" ULLS job - again with no batch
+ * buffer and, crucially, with no postamble semaphore wait or tail update -
+ * which lets the ring drain so the context can be switched off the hardware.
+ * The PM reference is then dropped. If the queue was not idle, the worker
+ * simply re-arms itself.
+ *
+ * Job state
+ * ---------
+ *
+ * The state above is communicated to the ring ops and GuC backend via
+ * @xe_sched_job.ulls, set under @xe_migrate.job_mutex:
+ *
+ * - %ULLS_NONE: job submitted outside of ULLS mode
+ * - %ULLS_ENTER: job that enters ULLS mode
+ * - %ULLS_ACTIVE: job submitted while in ULLS mode
+ * - %ULLS_EXIT: job that exits ULLS mode
+ */
+
/**
* struct xe_migrate - migrate context.
*/
--
2.34.1
^ permalink raw reply related [flat|nested] 54+ messages in thread* ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev9)
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (23 preceding siblings ...)
2026-09-25 4:53 ` [PATCH v7 24/24] drm/xe: Document ULLS for migration jobs Matthew Brost
@ 2026-09-25 5:02 ` Patchwork
2026-09-25 5:04 ` ✓ CI.KUnit: success " Patchwork
` (3 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Patchwork @ 2026-09-25 5:02 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
== Series Details ==
Series: CPU binds and ULLS on migration queue (rev9)
URL : https://patchwork.freedesktop.org/series/149888/
State : warning
== Summary ==
+ KERNEL=/kernel
+ git clone https://gitlab.freedesktop.org/drm/maintainer-tools mt
Cloning into 'mt'...
warning: redirecting to https://gitlab.freedesktop.org/drm/maintainer-tools.git/
+ git -C mt rev-list -n1 origin/master
6bef5ce5c74dc76b4ee69ab2832c0d1fd63c91c7
+ cd /kernel
+ git config --global --add safe.directory /kernel
+ git log -n1
commit 3a56a3540929eab627d2fd966cbc28e905285748
Author: Matthew Brost <matthew.brost@intel.com>
Date: Thu Sep 24 21:53:20 2026 -0700
drm/xe: Document ULLS for migration jobs
Add a kernel-doc DOC section at the top of xe_migrate.c describing the
Ultra Low Latency Submission (ULLS) scheme used for migration jobs.
Cover the motivation (removing the H2G / GuC / context switch latency
from the page fault and SVM prefetch critical paths), the platform
requirements, the LRC PPHWSP semaphore layout and its relationship to
the migration queue job count, the fixed ULLS job size and why it is
needed, the ring preamble / postamble emitted by the ring ops including
the in-ring tail update, the semaphore-only submission fast path in the
GuC backend, and the enter / delayed exit flow along with the ULLS job
flags.
Hook the new section into Documentation/gpu/xe/xe_migrate.rst.
Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Assisted-by: Github-Copilot:Claude-opus-5
+ /mt/dim checkpatch 999838292166407bbe911a41dee23112c49e9d44 drm-intel
abb2c1d4d172 drm/xe: reference VM from PT BOs
8317b77fbd55 drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns
71b0236c7bd6 drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper
979325388805 drm/xe: Decouple exec queue idle check from LRC
fad4f41c74dd drm/xe: Add job count to GuC exec queue snapshot
4bc4c9e6ddb8 drm/xe: Update xe_bo_put_deferred arguments to include writeback flag
ea63c158a98f drm/xe: Update scheduler job layer to support PT jobs
566709a74592 drm/xe: Add helpers to access PT ops
e9e844ec18ab drm/xe: Add struct xe_pt_job_ops
c537663ae172 drm/xe: Update GuC submission backend to run PT jobs
c1eea266904d drm/xe: Store level in struct xe_vm_pgtable_update
4038cd863e42 drm/xe: Don't use migrate exec queue for page fault binds
1c7090201a4d drm/xe: Enable CPU binds for jobs
6ab805b115c4 drm/xe: Remove unused arguments from xe_migrate_pt_update_ops
f62130f433de drm/xe: Make bind queues operate cross-tile
-:305: CHECK:MACRO_ARG_REUSE: Macro argument reuse '__i' - possible side-effects?
#305: FILE: drivers/gpu/drm/xe/xe_exec_queue.h:17:
+#define for_each_tlb_inval(__q, __i) \
+ for (__i = 0; __i < XE_EXEC_QUEUE_TLB_INVAL_COUNT; ++__i) \
+ for_each_if((__q)->tlb_inval[__i].dep_scheduler)
total: 0 errors, 0 warnings, 1 checks, 623 lines checked
728ac6f798d7 drm/xe: Add CPU bind layer
-:34: WARNING:FILE_PATH_CHANGES: added, moved or deleted file(s), does MAINTAINERS need updating?
#34:
new file mode 100644
total: 0 errors, 1 warnings, 0 checks, 2309 lines checked
65b042c18a22 drm/xe: Add device flag to enable PT mirroring across tiles
173669f99d1e drm/xe: Add ULLS support to LRC
dbe8d3073807 drm/xe: Add ULLS migration job support to migration layer
9eca7b36b46b drm/xe: Add ULLS migration job support to ring ops
482ce0e50999 drm/xe: Add ULLS migration job support to GuC submission
fa925b52fefe drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
16b5ad66b307 drm/xe: add migrate ULLS period configfs attribute
3a56a3540929 drm/xe: Document ULLS for migration jobs
^ permalink raw reply [flat|nested] 54+ messages in thread* ✓ CI.KUnit: success for CPU binds and ULLS on migration queue (rev9)
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (24 preceding siblings ...)
2026-09-25 5:02 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev9) Patchwork
@ 2026-09-25 5:04 ` Patchwork
2026-09-25 5:47 ` ✓ Xe.CI.BAT: " Patchwork
` (2 subsequent siblings)
28 siblings, 0 replies; 54+ messages in thread
From: Patchwork @ 2026-09-25 5:04 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
== Series Details ==
Series: CPU binds and ULLS on migration queue (rev9)
URL : https://patchwork.freedesktop.org/series/149888/
State : success
== Summary ==
+ trap cleanup EXIT
+ kunitconfigs=('/kernel/drivers/gpu/tests/.kunitconfig' '/kernel/drivers/gpu/drm/xe/.kunitconfig' '/kernel/drivers/gpu/drm/tests/.kunitconfig' '/kernel/drivers/gpu/drm/ttm/tests/.kunitconfig' '/kernel/drivers/dma-buf/.kunitconfig')
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/tests/.kunitconfig
[05:02:11] Configuring KUnit Kernel ...
Generating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[05:02:16] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[05:02:36] Starting KUnit Kernel (1/1)...
[05:02:36] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[05:02:36] ============= refcount_interrupt (4 subtests) ==============
[05:02:36] [PASSED] test_single_irq_change
[05:02:36] [PASSED] test_nested_irq_change
[05:02:36] [PASSED] test_multiple_irq_change
[05:02:36] [PASSED] test_irq_save
[05:02:36] =============== [PASSED] refcount_interrupt ================
[05:02:36] ================= gpu_buddy (14 subtests) ==================
[05:02:36] [PASSED] gpu_test_buddy_alloc_limit
[05:02:36] [PASSED] gpu_test_buddy_alloc_optimistic
[05:02:36] [PASSED] gpu_test_buddy_alloc_pessimistic
[05:02:36] [PASSED] gpu_test_buddy_alloc_pathological
[05:02:36] [PASSED] gpu_test_buddy_alloc_contiguous
[05:02:36] [PASSED] gpu_test_buddy_alloc_clear
[05:02:36] [PASSED] gpu_test_buddy_alloc_range
[05:02:36] [PASSED] gpu_test_buddy_alloc_range_bias
[05:02:37] [PASSED] gpu_test_buddy_fragmentation_performance
[05:02:38] [PASSED] gpu_test_buddy_dirty_tracker_performance
[05:02:38] [PASSED] gpu_test_buddy_alloc_exceeds_max_order
[05:02:38] [PASSED] gpu_test_buddy_offset_aligned_allocation
[05:02:38] [PASSED] gpu_test_buddy_subtree_offset_alignment_stress
[05:02:38] [PASSED] gpu_test_buddy_addr_to_block
[05:02:38] ==================== [PASSED] gpu_buddy ====================
[05:02:38] ============================================================
[05:02:38] Testing complete. Ran 18 tests: passed: 18
[05:02:38] Elapsed time: 26.589s total, 4.445s configuring, 20.276s building, 1.814s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/xe/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/xe/.kunitconfig
[05:02:38] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[05:02:40] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[05:03:13] Starting KUnit Kernel (1/1)...
[05:03:13] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[05:03:14] ============= refcount_interrupt (4 subtests) ==============
[05:03:14] [PASSED] test_single_irq_change
[05:03:14] [PASSED] test_nested_irq_change
[05:03:14] [PASSED] test_multiple_irq_change
[05:03:14] [PASSED] test_irq_save
[05:03:14] =============== [PASSED] refcount_interrupt ================
[05:03:14] ================== guc_buf (11 subtests) ===================
[05:03:14] [PASSED] test_smallest
[05:03:14] [PASSED] test_largest
[05:03:14] [PASSED] test_granular
[05:03:14] [PASSED] test_unique
[05:03:14] [PASSED] test_overlap
[05:03:14] [PASSED] test_reusable
[05:03:14] [PASSED] test_too_big
[05:03:14] [PASSED] test_flush
[05:03:14] [PASSED] test_lookup
[05:03:14] [PASSED] test_data
[05:03:14] [PASSED] test_class
[05:03:14] ===================== [PASSED] guc_buf =====================
[05:03:14] =================== guc_dbm (7 subtests) ===================
[05:03:14] [PASSED] test_empty
[05:03:14] [PASSED] test_default
[05:03:14] ======================== test_size ========================
[05:03:14] [PASSED] 4
[05:03:14] [PASSED] 8
[05:03:14] [PASSED] 32
[05:03:14] [PASSED] 256
[05:03:14] ==================== [PASSED] test_size ====================
[05:03:14] ======================= test_reuse ========================
[05:03:14] [PASSED] 4
[05:03:14] [PASSED] 8
[05:03:14] [PASSED] 32
[05:03:14] [PASSED] 256
[05:03:14] =================== [PASSED] test_reuse ====================
[05:03:14] =================== test_range_overlap ====================
[05:03:14] [PASSED] 4
[05:03:14] [PASSED] 8
[05:03:14] [PASSED] 32
[05:03:14] [PASSED] 256
[05:03:14] =============== [PASSED] test_range_overlap ================
[05:03:14] =================== test_range_compact ====================
[05:03:14] [PASSED] 4
[05:03:14] [PASSED] 8
[05:03:14] [PASSED] 32
[05:03:14] [PASSED] 256
[05:03:14] =============== [PASSED] test_range_compact ================
[05:03:14] ==================== test_range_spare =====================
[05:03:14] [PASSED] 4
[05:03:14] [PASSED] 8
[05:03:14] [PASSED] 32
[05:03:14] [PASSED] 256
[05:03:14] ================ [PASSED] test_range_spare =================
[05:03:14] ===================== [PASSED] guc_dbm =====================
[05:03:14] =================== guc_idm (6 subtests) ===================
[05:03:14] [PASSED] bad_init
[05:03:14] [PASSED] no_init
[05:03:14] [PASSED] init_fini
[05:03:14] [PASSED] check_used
[05:03:14] [PASSED] check_quota
[05:03:14] [PASSED] check_all
[05:03:14] ===================== [PASSED] guc_idm =====================
[05:03:14] ============== guc_klv_helpers (13 subtests) ===============
[05:03:14] [PASSED] test_count
[05:03:14] [PASSED] test_encode_u32
[05:03:14] [PASSED] test_encode_u64
[05:03:14] [PASSED] test_encode_string
[05:03:14] [PASSED] test_encode_object_raw
[05:03:14] [PASSED] test_encode_object_klv
[05:03:14] [PASSED] test_encode_object_nested
[05:03:14] [PASSED] test_encode_object_basic
[05:03:14] ===================== test_decode_u16 =====================
[05:03:14] [PASSED] empty
[05:03:14] [PASSED] valid
[05:03:14] [PASSED] max
[05:03:14] [PASSED] big
[05:03:14] [PASSED] long
[05:03:14] ================= [PASSED] test_decode_u16 =================
[05:03:14] ===================== test_decode_u32 =====================
[05:03:14] [PASSED] empty
[05:03:14] [PASSED] valid
[05:03:14] [PASSED] max
[05:03:14] [PASSED] long
[05:03:14] ================= [PASSED] test_decode_u32 =================
[05:03:14] ===================== test_decode_u64 =====================
[05:03:14] [PASSED] empty
[05:03:14] [PASSED] short
[05:03:14] [PASSED] valid
[05:03:14] [PASSED] max
[05:03:14] [PASSED] long
[05:03:14] ================= [PASSED] test_decode_u64 =================
[05:03:14] ===================== test_decode_str =====================
[05:03:14] [PASSED] empty
[05:03:14] [PASSED] one
[05:03:14] [PASSED] one_trash
[05:03:14] [PASSED] two
[05:03:14] [PASSED] two_trash
[05:03:14] [PASSED] three
[05:03:14] [PASSED] dword
[05:03:14] [PASSED] four
[05:03:14] [PASSED] five
[05:03:14] [PASSED] six
[05:03:14] [PASSED] seven
[05:03:14] [PASSED] qword
[05:03:14] ================= [PASSED] test_decode_str =================
[05:03:14] [PASSED] test_print
[05:03:14] ================= [PASSED] guc_klv_helpers =================
[05:03:14] =================== xe_log (4 subtests) ====================
[05:03:14] [PASSED] demo_cper
[05:03:14] [PASSED] demo_dmesg
[05:03:14] ======================= test_dmesg ========================
[05:03:14] [PASSED] test_fatal
[05:03:14] [PASSED] test_fatal_tile
[05:03:14] [PASSED] test_fatal_gt
[05:03:14] [PASSED] test_fatal_comp
[05:03:14] [PASSED] test_fatal_comp_tile
[05:03:14] [PASSED] test_fatal_comp_gt
[05:03:14] [PASSED] test_fatal_all
[05:03:14] [PASSED] test_recoverable
[05:03:14] [PASSED] test_recoverable_tile
[05:03:14] [PASSED] test_recoverable_gt
[05:03:14] [PASSED] test_recoverable_comp
[05:03:14] [PASSED] test_recoverable_comp_tile
[05:03:14] [PASSED] test_recoverable_comp_gt
[05:03:14] [PASSED] test_recoverable_all
[05:03:14] [PASSED] test_info
[05:03:14] [PASSED] test_info_tile
[05:03:14] [PASSED] test_info_gt
[05:03:14] [PASSED] test_info_err
[05:03:14] [PASSED] test_info_comp
[05:03:14] [PASSED] test_info_comp_tile
[05:03:14] [PASSED] test_info_comp_gt
[05:03:14] [PASSED] test_info_all
[05:03:14] [PASSED] test_hw_fatal
[05:03:14] [PASSED] test_hw_recoverable
[05:03:14] [PASSED] test_hw_corrected
[05:03:14] [PASSED] test_hw_informational
[05:03:14] =================== [PASSED] test_dmesg ====================
[05:03:14] ====================== test_invalid =======================
[05:03:14] [SKIPPED] no-component no-location no-warn (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] reserved location (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] unknown location (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] nonzero-device-id location (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] invalid-tile-id location (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] invalid-gt-id location (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] unknown component class (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] unknown system component (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] unknown hardware component (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] [SKIPPED] unknown component and location (requires CONFIG_DRM_XE_DEBUG)
[05:03:14] ================== [SKIPPED] test_invalid ==================
[05:03:14] ===================== [PASSED] xe_log ======================
[05:03:14] ================== no_relay (3 subtests) ===================
[05:03:14] [PASSED] xe_drops_guc2pf_if_not_ready
[05:03:14] [PASSED] xe_drops_guc2vf_if_not_ready
[05:03:14] [PASSED] xe_rejects_send_if_not_ready
[05:03:14] ==================== [PASSED] no_relay =====================
[05:03:14] ================== pf_relay (14 subtests) ==================
[05:03:14] [PASSED] pf_rejects_guc2pf_too_short
[05:03:14] [PASSED] pf_rejects_guc2pf_too_long
[05:03:14] [PASSED] pf_rejects_guc2pf_no_payload
[05:03:14] [PASSED] pf_fails_no_payload
[05:03:14] [PASSED] pf_fails_bad_origin
[05:03:14] [PASSED] pf_fails_bad_type
[05:03:14] [PASSED] pf_txn_reports_error
[05:03:14] [PASSED] pf_txn_sends_pf2guc
[05:03:14] [PASSED] pf_sends_pf2guc
[05:03:14] [SKIPPED] pf_loopback_nop (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[05:03:14] [SKIPPED] pf_loopback_echo (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[05:03:14] [SKIPPED] pf_loopback_fail (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[05:03:14] [SKIPPED] pf_loopback_busy (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[05:03:14] [SKIPPED] pf_loopback_retry (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[05:03:14] ==================== [PASSED] pf_relay =====================
[05:03:14] ================== vf_relay (3 subtests) ===================
[05:03:14] [PASSED] vf_rejects_guc2vf_too_short
[05:03:14] [PASSED] vf_rejects_guc2vf_too_long
[05:03:14] [PASSED] vf_rejects_guc2vf_no_payload
[05:03:14] ==================== [PASSED] vf_relay =====================
[05:03:14] ================ pf_gt_config (9 subtests) =================
[05:03:14] [PASSED] fair_contexts_1vf
[05:03:14] [PASSED] fair_doorbells_1vf
[05:03:14] [PASSED] fair_ggtt_1vf
[05:03:14] ====================== fair_vram_1vf ======================
[05:03:14] [PASSED] 3.50 GiB
[05:03:14] [PASSED] 11.5 GiB
[05:03:14] [PASSED] 15.5 GiB
[05:03:14] [PASSED] 31.5 GiB
[05:03:14] [PASSED] 63.5 GiB
[05:03:14] [PASSED] 1.91 GiB
[05:03:14] ================== [PASSED] fair_vram_1vf ==================
[05:03:14] ================ fair_vram_1vf_admin_only =================
[05:03:14] [PASSED] 3.50 GiB
[05:03:14] [PASSED] 11.5 GiB
[05:03:14] [PASSED] 15.5 GiB
[05:03:14] [PASSED] 31.5 GiB
[05:03:14] [PASSED] 63.5 GiB
[05:03:14] [PASSED] 1.91 GiB
[05:03:14] ============ [PASSED] fair_vram_1vf_admin_only =============
[05:03:14] ====================== fair_contexts ======================
[05:03:14] [PASSED] 1 VF
[05:03:14] [PASSED] 2 VFs
[05:03:14] [PASSED] 3 VFs
[05:03:14] [PASSED] 4 VFs
[05:03:14] [PASSED] 5 VFs
[05:03:14] [PASSED] 6 VFs
[05:03:14] [PASSED] 7 VFs
[05:03:14] [PASSED] 8 VFs
[05:03:14] [PASSED] 9 VFs
[05:03:14] [PASSED] 10 VFs
[05:03:14] [PASSED] 11 VFs
[05:03:14] [PASSED] 12 VFs
[05:03:14] [PASSED] 13 VFs
[05:03:14] [PASSED] 14 VFs
[05:03:14] [PASSED] 15 VFs
[05:03:14] [PASSED] 16 VFs
[05:03:14] [PASSED] 17 VFs
[05:03:14] [PASSED] 18 VFs
[05:03:14] [PASSED] 19 VFs
[05:03:14] [PASSED] 20 VFs
[05:03:14] [PASSED] 21 VFs
[05:03:14] [PASSED] 22 VFs
[05:03:14] [PASSED] 23 VFs
[05:03:14] [PASSED] 24 VFs
[05:03:14] [PASSED] 25 VFs
[05:03:14] [PASSED] 26 VFs
[05:03:14] [PASSED] 27 VFs
[05:03:14] [PASSED] 28 VFs
[05:03:14] [PASSED] 29 VFs
[05:03:14] [PASSED] 30 VFs
[05:03:14] [PASSED] 31 VFs
[05:03:14] [PASSED] 32 VFs
[05:03:14] [PASSED] 33 VFs
[05:03:14] [PASSED] 34 VFs
[05:03:14] [PASSED] 35 VFs
[05:03:14] [PASSED] 36 VFs
[05:03:14] [PASSED] 37 VFs
[05:03:14] [PASSED] 38 VFs
[05:03:14] [PASSED] 39 VFs
[05:03:14] [PASSED] 40 VFs
[05:03:14] [PASSED] 41 VFs
[05:03:14] [PASSED] 42 VFs
[05:03:14] [PASSED] 43 VFs
[05:03:14] [PASSED] 44 VFs
[05:03:14] [PASSED] 45 VFs
[05:03:14] [PASSED] 46 VFs
[05:03:14] [PASSED] 47 VFs
[05:03:14] [PASSED] 48 VFs
[05:03:14] [PASSED] 49 VFs
[05:03:14] [PASSED] 50 VFs
[05:03:14] [PASSED] 51 VFs
[05:03:14] [PASSED] 52 VFs
[05:03:14] [PASSED] 53 VFs
[05:03:14] [PASSED] 54 VFs
[05:03:14] [PASSED] 55 VFs
[05:03:14] [PASSED] 56 VFs
[05:03:14] [PASSED] 57 VFs
[05:03:14] [PASSED] 58 VFs
[05:03:14] [PASSED] 59 VFs
[05:03:14] [PASSED] 60 VFs
[05:03:14] [PASSED] 61 VFs
[05:03:14] [PASSED] 62 VFs
[05:03:14] [PASSED] 63 VFs
[05:03:14] ================== [PASSED] fair_contexts ==================
[05:03:14] ===================== fair_doorbells ======================
[05:03:14] [PASSED] 1 VF
[05:03:14] [PASSED] 2 VFs
[05:03:14] [PASSED] 3 VFs
[05:03:14] [PASSED] 4 VFs
[05:03:14] [PASSED] 5 VFs
[05:03:14] [PASSED] 6 VFs
[05:03:14] [PASSED] 7 VFs
[05:03:14] [PASSED] 8 VFs
[05:03:14] [PASSED] 9 VFs
[05:03:14] [PASSED] 10 VFs
[05:03:14] [PASSED] 11 VFs
[05:03:14] [PASSED] 12 VFs
[05:03:14] [PASSED] 13 VFs
[05:03:14] [PASSED] 14 VFs
[05:03:14] [PASSED] 15 VFs
[05:03:14] [PASSED] 16 VFs
[05:03:14] [PASSED] 17 VFs
[05:03:14] [PASSED] 18 VFs
[05:03:14] [PASSED] 19 VFs
[05:03:14] [PASSED] 20 VFs
[05:03:14] [PASSED] 21 VFs
[05:03:14] [PASSED] 22 VFs
[05:03:14] [PASSED] 23 VFs
[05:03:14] [PASSED] 24 VFs
[05:03:14] [PASSED] 25 VFs
[05:03:14] [PASSED] 26 VFs
[05:03:14] [PASSED] 27 VFs
[05:03:14] [PASSED] 28 VFs
[05:03:14] [PASSED] 29 VFs
[05:03:14] [PASSED] 30 VFs
[05:03:14] [PASSED] 31 VFs
[05:03:14] [PASSED] 32 VFs
[05:03:14] [PASSED] 33 VFs
[05:03:14] [PASSED] 34 VFs
[05:03:14] [PASSED] 35 VFs
[05:03:14] [PASSED] 36 VFs
[05:03:14] [PASSED] 37 VFs
[05:03:14] [PASSED] 38 VFs
[05:03:14] [PASSED] 39 VFs
[05:03:14] [PASSED] 40 VFs
[05:03:14] [PASSED] 41 VFs
[05:03:14] [PASSED] 42 VFs
[05:03:14] [PASSED] 43 VFs
[05:03:14] [PASSED] 44 VFs
[05:03:14] [PASSED] 45 VFs
[05:03:14] [PASSED] 46 VFs
[05:03:14] [PASSED] 47 VFs
[05:03:14] [PASSED] 48 VFs
[05:03:14] [PASSED] 49 VFs
[05:03:14] [PASSED] 50 VFs
[05:03:14] [PASSED] 51 VFs
[05:03:14] [PASSED] 52 VFs
[05:03:14] [PASSED] 53 VFs
[05:03:14] [PASSED] 54 VFs
[05:03:14] [PASSED] 55 VFs
[05:03:14] [PASSED] 56 VFs
[05:03:14] [PASSED] 57 VFs
[05:03:14] [PASSED] 58 VFs
[05:03:14] [PASSED] 59 VFs
[05:03:14] [PASSED] 60 VFs
[05:03:14] [PASSED] 61 VFs
[05:03:14] [PASSED] 62 VFs
[05:03:14] [PASSED] 63 VFs
[05:03:14] ================= [PASSED] fair_doorbells ==================
[05:03:14] ======================== fair_ggtt ========================
[05:03:14] [PASSED] 1 VF
[05:03:14] [PASSED] 2 VFs
[05:03:14] [PASSED] 3 VFs
[05:03:14] [PASSED] 4 VFs
[05:03:14] [PASSED] 5 VFs
[05:03:14] [PASSED] 6 VFs
[05:03:14] [PASSED] 7 VFs
[05:03:14] [PASSED] 8 VFs
[05:03:14] [PASSED] 9 VFs
[05:03:14] [PASSED] 10 VFs
[05:03:14] [PASSED] 11 VFs
[05:03:14] [PASSED] 12 VFs
[05:03:14] [PASSED] 13 VFs
[05:03:14] [PASSED] 14 VFs
[05:03:14] [PASSED] 15 VFs
[05:03:14] [PASSED] 16 VFs
[05:03:14] [PASSED] 17 VFs
[05:03:14] [PASSED] 18 VFs
[05:03:14] [PASSED] 19 VFs
[05:03:14] [PASSED] 20 VFs
[05:03:14] [PASSED] 21 VFs
[05:03:14] [PASSED] 22 VFs
[05:03:14] [PASSED] 23 VFs
[05:03:14] [PASSED] 24 VFs
[05:03:14] [PASSED] 25 VFs
[05:03:14] [PASSED] 26 VFs
[05:03:14] [PASSED] 27 VFs
[05:03:14] [PASSED] 28 VFs
[05:03:14] [PASSED] 29 VFs
[05:03:14] [PASSED] 30 VFs
[05:03:14] [PASSED] 31 VFs
[05:03:14] [PASSED] 32 VFs
[05:03:14] [PASSED] 33 VFs
[05:03:14] [PASSED] 34 VFs
[05:03:14] [PASSED] 35 VFs
[05:03:14] [PASSED] 36 VFs
[05:03:14] [PASSED] 37 VFs
[05:03:14] [PASSED] 38 VFs
[05:03:14] [PASSED] 39 VFs
[05:03:14] [PASSED] 40 VFs
[05:03:14] [PASSED] 41 VFs
[05:03:14] [PASSED] 42 VFs
[05:03:14] [PASSED] 43 VFs
[05:03:14] [PASSED] 44 VFs
[05:03:14] [PASSED] 45 VFs
[05:03:14] [PASSED] 46 VFs
[05:03:14] [PASSED] 47 VFs
[05:03:14] [PASSED] 48 VFs
[05:03:14] [PASSED] 49 VFs
[05:03:14] [PASSED] 50 VFs
[05:03:14] [PASSED] 51 VFs
[05:03:14] [PASSED] 52 VFs
[05:03:14] [PASSED] 53 VFs
[05:03:14] [PASSED] 54 VFs
[05:03:14] [PASSED] 55 VFs
[05:03:14] [PASSED] 56 VFs
[05:03:14] [PASSED] 57 VFs
[05:03:14] [PASSED] 58 VFs
[05:03:14] [PASSED] 59 VFs
[05:03:14] [PASSED] 60 VFs
[05:03:14] [PASSED] 61 VFs
[05:03:14] [PASSED] 62 VFs
[05:03:14] [PASSED] 63 VFs
[05:03:14] ==================== [PASSED] fair_ggtt ====================
[05:03:14] ======================== fair_vram ========================
[05:03:14] [PASSED] 1 VF
[05:03:14] [PASSED] 2 VFs
[05:03:14] [PASSED] 3 VFs
[05:03:14] [PASSED] 4 VFs
[05:03:14] [PASSED] 5 VFs
[05:03:14] [PASSED] 6 VFs
[05:03:14] [PASSED] 7 VFs
[05:03:14] [PASSED] 8 VFs
[05:03:14] [PASSED] 9 VFs
[05:03:14] [PASSED] 10 VFs
[05:03:14] [PASSED] 11 VFs
[05:03:14] [PASSED] 12 VFs
[05:03:14] [PASSED] 13 VFs
[05:03:14] [PASSED] 14 VFs
[05:03:14] [PASSED] 15 VFs
[05:03:14] [PASSED] 16 VFs
[05:03:14] [PASSED] 17 VFs
[05:03:14] [PASSED] 18 VFs
[05:03:14] [PASSED] 19 VFs
[05:03:14] [PASSED] 20 VFs
[05:03:14] [PASSED] 21 VFs
[05:03:14] [PASSED] 22 VFs
[05:03:14] [PASSED] 23 VFs
[05:03:14] [PASSED] 24 VFs
[05:03:14] [PASSED] 25 VFs
[05:03:14] [PASSED] 26 VFs
[05:03:14] [PASSED] 27 VFs
[05:03:14] [PASSED] 28 VFs
[05:03:14] [PASSED] 29 VFs
[05:03:14] [PASSED] 30 VFs
[05:03:14] [PASSED] 31 VFs
[05:03:14] [PASSED] 32 VFs
[05:03:14] [PASSED] 33 VFs
[05:03:14] [PASSED] 34 VFs
[05:03:14] [PASSED] 35 VFs
[05:03:14] [PASSED] 36 VFs
[05:03:14] [PASSED] 37 VFs
[05:03:14] [PASSED] 38 VFs
[05:03:14] [PASSED] 39 VFs
[05:03:14] [PASSED] 40 VFs
[05:03:14] [PASSED] 41 VFs
[05:03:14] [PASSED] 42 VFs
[05:03:14] [PASSED] 43 VFs
[05:03:14] [PASSED] 44 VFs
[05:03:14] [PASSED] 45 VFs
[05:03:14] [PASSED] 46 VFs
[05:03:14] [PASSED] 47 VFs
[05:03:14] [PASSED] 48 VFs
[05:03:14] [PASSED] 49 VFs
[05:03:14] [PASSED] 50 VFs
[05:03:14] [PASSED] 51 VFs
[05:03:14] [PASSED] 52 VFs
[05:03:14] [PASSED] 53 VFs
[05:03:14] [PASSED] 54 VFs
[05:03:14] [PASSED] 55 VFs
[05:03:14] [PASSED] 56 VFs
[05:03:14] [PASSED] 57 VFs
[05:03:14] [PASSED] 58 VFs
[05:03:14] [PASSED] 59 VFs
[05:03:14] [PASSED] 60 VFs
[05:03:14] [PASSED] 61 VFs
[05:03:14] [PASSED] 62 VFs
[05:03:14] [PASSED] 63 VFs
[05:03:14] ==================== [PASSED] fair_vram ====================
[05:03:14] ================== [PASSED] pf_gt_config ===================
[05:03:14] ===================== lmtt (1 subtest) =====================
[05:03:14] ======================== test_ops =========================
[05:03:14] [PASSED] 2-level
[05:03:14] [PASSED] multi-level
[05:03:14] ==================== [PASSED] test_ops =====================
[05:03:14] ====================== [PASSED] lmtt =======================
[05:03:14] ================= sriov_packet (1 subtest) =================
[05:03:14] [PASSED] test_descriptor_init
[05:03:14] ================== [PASSED] sriov_packet ===================
[05:03:14] ================= pf_service (11 subtests) =================
[05:03:14] [PASSED] pf_negotiate_any
[05:03:14] [PASSED] pf_negotiate_base_match
[05:03:14] [PASSED] pf_negotiate_base_newer
[05:03:14] [PASSED] pf_negotiate_base_next
[05:03:14] [SKIPPED] pf_negotiate_base_older (no older minor)
[05:03:14] [PASSED] pf_negotiate_base_prev
[05:03:14] [PASSED] pf_negotiate_latest_match
[05:03:14] [PASSED] pf_negotiate_latest_newer
[05:03:14] [PASSED] pf_negotiate_latest_next
[05:03:14] [SKIPPED] pf_negotiate_latest_older (no older minor)
[05:03:14] [SKIPPED] pf_negotiate_latest_prev (no prev major)
[05:03:14] =================== [PASSED] pf_service ====================
[05:03:14] ================= xe_guc_g2g (2 subtests) ==================
[05:03:14] ============== xe_live_guc_g2g_kunit_default ==============
[05:03:14] ========= [SKIPPED] xe_live_guc_g2g_kunit_default ==========
[05:03:14] ============== xe_live_guc_g2g_kunit_allmem ===============
[05:03:14] ========== [SKIPPED] xe_live_guc_g2g_kunit_allmem ==========
[05:03:14] =================== [SKIPPED] xe_guc_g2g ===================
[05:03:14] =================== xe_mocs (2 subtests) ===================
[05:03:14] ================ xe_live_mocs_kernel_kunit ================
[05:03:14] =========== [SKIPPED] xe_live_mocs_kernel_kunit ============
[05:03:14] ================ xe_live_mocs_reset_kunit =================
[05:03:14] ============ [SKIPPED] xe_live_mocs_reset_kunit ============
[05:03:14] ==================== [SKIPPED] xe_mocs =====================
[05:03:14] ================= xe_migrate (2 subtests) ==================
[05:03:14] ================= xe_migrate_sanity_kunit =================
[05:03:14] ============ [SKIPPED] xe_migrate_sanity_kunit =============
[05:03:14] ================== xe_validate_ccs_kunit ==================
[05:03:14] ============= [SKIPPED] xe_validate_ccs_kunit ==============
[05:03:14] =================== [SKIPPED] xe_migrate ===================
[05:03:14] ================== xe_dma_buf (1 subtest) ==================
[05:03:14] ==================== xe_dma_buf_kunit =====================
[05:03:14] ================ [SKIPPED] xe_dma_buf_kunit ================
[05:03:14] =================== [SKIPPED] xe_dma_buf ===================
[05:03:14] ================= xe_bo_shrink (1 subtest) =================
[05:03:14] =================== xe_bo_shrink_kunit ====================
[05:03:14] =============== [SKIPPED] xe_bo_shrink_kunit ===============
[05:03:14] ================== [SKIPPED] xe_bo_shrink ==================
[05:03:14] ==================== xe_bo (2 subtests) ====================
[05:03:14] ================== xe_ccs_migrate_kunit ===================
[05:03:14] ============== [SKIPPED] xe_ccs_migrate_kunit ==============
[05:03:14] ==================== xe_bo_evict_kunit ====================
[05:03:14] =============== [SKIPPED] xe_bo_evict_kunit ================
[05:03:14] ===================== [SKIPPED] xe_bo ======================
[05:03:14] =================== xe_any (9 subtests) ====================
[05:03:14] [PASSED] test_to_xe
[05:03:14] [PASSED] test_to_dev
[05:03:14] [PASSED] test_to_pdev
[05:03:14] [PASSED] test_to_drm
[05:03:14] [PASSED] test_if_pdev
[05:03:14] [PASSED] test_if_xe
[05:03:14] [PASSED] test_if_tile
[05:03:14] [PASSED] test_if_gt
[05:03:14] [PASSED] test_to_id
[05:03:14] ===================== [PASSED] xe_any ======================
[05:03:14] ==================== args (13 subtests) ====================
[05:03:14] [PASSED] count_args_test
[05:03:14] [PASSED] call_args_example
[05:03:14] [PASSED] call_args_test
[05:03:14] [PASSED] drop_first_arg_example
[05:03:14] [PASSED] drop_first_arg_test
[05:03:14] [PASSED] first_arg_example
[05:03:14] [PASSED] first_arg_test
[05:03:14] [PASSED] last_arg_example
[05:03:14] [PASSED] last_arg_test
[05:03:14] [PASSED] pick_arg_example
[05:03:14] [PASSED] if_args_example
[05:03:14] [PASSED] if_args_test
[05:03:14] [PASSED] sep_comma_example
[05:03:14] ====================== [PASSED] args =======================
[05:03:14] =================== xe_pci (3 subtests) ====================
[05:03:14] ==================== check_graphics_ip ====================
[05:03:14] [PASSED] 12.00 Xe_LP
[05:03:14] [PASSED] 12.10 Xe_LP+
[05:03:14] [PASSED] 12.55 Xe_HPG
[05:03:14] [PASSED] 12.60 Xe_HPC
[05:03:14] [PASSED] 12.70 Xe_LPG
[05:03:14] [PASSED] 12.71 Xe_LPG
[05:03:14] [PASSED] 12.74 Xe_LPG+
[05:03:14] [PASSED] 20.01 Xe2_HPG
[05:03:14] [PASSED] 20.02 Xe2_HPG
[05:03:14] [PASSED] 20.04 Xe2_LPG
[05:03:14] [PASSED] 30.00 Xe3_LPG
[05:03:14] [PASSED] 30.01 Xe3_LPG
[05:03:14] [PASSED] 30.03 Xe3_LPG
[05:03:14] [PASSED] 30.04 Xe3_LPG
[05:03:14] [PASSED] 30.05 Xe3_LPG
[05:03:14] [PASSED] 35.10 Xe3p_LPG
[05:03:14] [PASSED] 35.11 Xe3p_XPC
[05:03:14] ================ [PASSED] check_graphics_ip ================
[05:03:14] ===================== check_media_ip ======================
[05:03:14] [PASSED] 12.00 Xe_M
[05:03:14] [PASSED] 12.55 Xe_HPM
[05:03:14] [PASSED] 13.00 Xe_LPM+
[05:03:14] [PASSED] 13.01 Xe2_HPM
[05:03:14] [PASSED] 20.00 Xe2_LPM
[05:03:14] [PASSED] 30.00 Xe3_LPM
[05:03:14] [PASSED] 30.02 Xe3_LPM
[05:03:14] [PASSED] 35.00 Xe3p_LPM
[05:03:14] [PASSED] 35.03 Xe3p_HPM
[05:03:14] ================= [PASSED] check_media_ip ==================
[05:03:14] =================== check_platform_desc ===================
[05:03:14] [PASSED] 0x9A60 (TIGERLAKE)
[05:03:14] [PASSED] 0x9A68 (TIGERLAKE)
[05:03:14] [PASSED] 0x9A70 (TIGERLAKE)
[05:03:14] [PASSED] 0x9A40 (TIGERLAKE)
[05:03:14] [PASSED] 0x9A49 (TIGERLAKE)
[05:03:14] [PASSED] 0x9A59 (TIGERLAKE)
[05:03:14] [PASSED] 0x9A78 (TIGERLAKE)
[05:03:14] [PASSED] 0x9AC0 (TIGERLAKE)
[05:03:14] [PASSED] 0x9AC9 (TIGERLAKE)
[05:03:14] [PASSED] 0x9AD9 (TIGERLAKE)
[05:03:14] [PASSED] 0x9AF8 (TIGERLAKE)
[05:03:14] [PASSED] 0x4C80 (ROCKETLAKE)
[05:03:14] [PASSED] 0x4C8A (ROCKETLAKE)
[05:03:14] [PASSED] 0x4C8B (ROCKETLAKE)
[05:03:14] [PASSED] 0x4C8C (ROCKETLAKE)
[05:03:14] [PASSED] 0x4C90 (ROCKETLAKE)
[05:03:14] [PASSED] 0x4C9A (ROCKETLAKE)
[05:03:14] [PASSED] 0x4680 (ALDERLAKE_S)
[05:03:14] [PASSED] 0x4682 (ALDERLAKE_S)
[05:03:14] [PASSED] 0x4688 (ALDERLAKE_S)
[05:03:14] [PASSED] 0x468A (ALDERLAKE_S)
[05:03:14] [PASSED] 0x468B (ALDERLAKE_S)
[05:03:14] [PASSED] 0x4690 (ALDERLAKE_S)
[05:03:14] [PASSED] 0x4692 (ALDERLAKE_S)
[05:03:14] [PASSED] 0x4693 (ALDERLAKE_S)
[05:03:14] [PASSED] 0x46A0 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46A1 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46A2 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46A3 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46A6 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46A8 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46AA (ALDERLAKE_P)
[05:03:14] [PASSED] 0x462A (ALDERLAKE_P)
[05:03:14] [PASSED] 0x4626 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x4628 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46B0 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46B1 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46B2 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46B3 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46C0 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46C1 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46C2 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46C3 (ALDERLAKE_P)
[05:03:14] [PASSED] 0x46D0 (ALDERLAKE_N)
[05:03:14] [PASSED] 0x46D1 (ALDERLAKE_N)
[05:03:14] [PASSED] 0x46D2 (ALDERLAKE_N)
[05:03:14] [PASSED] 0x46D3 (ALDERLAKE_N)
[05:03:14] [PASSED] 0x46D4 (ALDERLAKE_N)
[05:03:14] [PASSED] 0xA721 (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7A1 (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7A9 (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7AC (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7AD (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA720 (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7A0 (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7A8 (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7AA (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA7AB (ALDERLAKE_P)
[05:03:14] [PASSED] 0xA780 (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA781 (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA782 (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA783 (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA788 (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA789 (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA78A (ALDERLAKE_S)
[05:03:14] [PASSED] 0xA78B (ALDERLAKE_S)
[05:03:14] [PASSED] 0x4905 (DG1)
[05:03:14] [PASSED] 0x4906 (DG1)
[05:03:14] [PASSED] 0x4907 (DG1)
[05:03:14] [PASSED] 0x4908 (DG1)
[05:03:14] [PASSED] 0x4909 (DG1)
[05:03:14] [PASSED] 0x56C0 (DG2)
[05:03:14] [PASSED] 0x56C2 (DG2)
[05:03:14] [PASSED] 0x56C1 (DG2)
[05:03:14] [PASSED] 0x7D51 (METEORLAKE)
[05:03:14] [PASSED] 0x7DD1 (METEORLAKE)
[05:03:14] [PASSED] 0x7D41 (METEORLAKE)
[05:03:14] [PASSED] 0x7D67 (METEORLAKE)
[05:03:14] [PASSED] 0xB640 (METEORLAKE)
[05:03:14] [PASSED] 0x56A0 (DG2)
[05:03:14] [PASSED] 0x56A1 (DG2)
[05:03:14] [PASSED] 0x56A2 (DG2)
[05:03:14] [PASSED] 0x56BE (DG2)
[05:03:14] [PASSED] 0x56BF (DG2)
[05:03:14] [PASSED] 0x5690 (DG2)
[05:03:14] [PASSED] 0x5691 (DG2)
[05:03:14] [PASSED] 0x5692 (DG2)
[05:03:14] [PASSED] 0x56A5 (DG2)
[05:03:14] [PASSED] 0x56A6 (DG2)
[05:03:14] [PASSED] 0x56B0 (DG2)
[05:03:14] [PASSED] 0x56B1 (DG2)
[05:03:14] [PASSED] 0x56BA (DG2)
[05:03:14] [PASSED] 0x56BB (DG2)
[05:03:14] [PASSED] 0x56BC (DG2)
[05:03:14] [PASSED] 0x56BD (DG2)
[05:03:14] [PASSED] 0x5693 (DG2)
[05:03:14] [PASSED] 0x5694 (DG2)
[05:03:14] [PASSED] 0x5695 (DG2)
[05:03:14] [PASSED] 0x56A3 (DG2)
[05:03:14] [PASSED] 0x56A4 (DG2)
[05:03:14] [PASSED] 0x56B2 (DG2)
[05:03:14] [PASSED] 0x56B3 (DG2)
[05:03:14] [PASSED] 0x5696 (DG2)
[05:03:14] [PASSED] 0x5697 (DG2)
[05:03:14] [PASSED] 0xB69 (PVC)
[05:03:14] [PASSED] 0xB6E (PVC)
[05:03:14] [PASSED] 0xBD4 (PVC)
[05:03:14] [PASSED] 0xBD5 (PVC)
[05:03:14] [PASSED] 0xBD6 (PVC)
[05:03:14] [PASSED] 0xBD7 (PVC)
[05:03:14] [PASSED] 0xBD8 (PVC)
[05:03:14] [PASSED] 0xBD9 (PVC)
[05:03:14] [PASSED] 0xBDA (PVC)
[05:03:14] [PASSED] 0xBDB (PVC)
[05:03:14] [PASSED] 0xBE0 (PVC)
[05:03:14] [PASSED] 0xBE1 (PVC)
[05:03:14] [PASSED] 0xBE5 (PVC)
[05:03:14] [PASSED] 0x7D40 (METEORLAKE)
[05:03:14] [PASSED] 0x7D45 (METEORLAKE)
[05:03:14] [PASSED] 0x7D55 (METEORLAKE)
[05:03:14] [PASSED] 0x7D60 (METEORLAKE)
[05:03:14] [PASSED] 0x7DD5 (METEORLAKE)
[05:03:14] [PASSED] 0x6420 (LUNARLAKE)
[05:03:14] [PASSED] 0x64A0 (LUNARLAKE)
[05:03:14] [PASSED] 0x64B0 (LUNARLAKE)
[05:03:14] [PASSED] 0xE202 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE209 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE20B (BATTLEMAGE)
[05:03:14] [PASSED] 0xE20C (BATTLEMAGE)
[05:03:14] [PASSED] 0xE20D (BATTLEMAGE)
[05:03:14] [PASSED] 0xE210 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE211 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE212 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE216 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE220 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE221 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE222 (BATTLEMAGE)
[05:03:14] [PASSED] 0xE223 (BATTLEMAGE)
[05:03:14] [PASSED] 0xB080 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB081 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB082 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB083 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB084 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB085 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB086 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB087 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB08F (PANTHERLAKE)
[05:03:14] [PASSED] 0xB090 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB0A0 (PANTHERLAKE)
[05:03:14] [PASSED] 0xB0B0 (PANTHERLAKE)
[05:03:14] [PASSED] 0xFD80 (PANTHERLAKE)
[05:03:14] [PASSED] 0xFD81 (PANTHERLAKE)
[05:03:14] [PASSED] 0xD740 (NOVALAKE_S)
[05:03:14] [PASSED] 0xD741 (NOVALAKE_S)
[05:03:14] [PASSED] 0xD742 (NOVALAKE_S)
[05:03:14] [PASSED] 0xD743 (NOVALAKE_S)
[05:03:14] [PASSED] 0xD745 (NOVALAKE_S)
[05:03:14] [PASSED] 0xD74A (NOVALAKE_S)
[05:03:14] [PASSED] 0xD74B (NOVALAKE_S)
[05:03:14] [PASSED] 0x674C (CRESCENTISLAND)
[05:03:14] [PASSED] 0x674D (CRESCENTISLAND)
[05:03:14] [PASSED] 0x674E (CRESCENTISLAND)
[05:03:14] [PASSED] 0x674F (CRESCENTISLAND)
[05:03:14] [PASSED] 0x6750 (CRESCENTISLAND)
[05:03:14] [PASSED] 0xD750 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD751 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD752 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD753 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD754 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD755 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD756 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD757 (NOVALAKE_P)
[05:03:14] [PASSED] 0xD75F (NOVALAKE_P)
[05:03:14] =============== [PASSED] check_platform_desc ===============
[05:03:14] ===================== [PASSED] xe_pci ======================
[05:03:14] ============= xe_rtp_tables_test (5 subtests) ==============
[05:03:14] ================== xe_rtp_table_gt_test ===================
[05:03:14] [PASSED] gt_was/14011060649
[05:03:14] [PASSED] gt_was/14011059788
[05:03:14] [PASSED] gt_was/14015795083
[05:03:14] [PASSED] gt_was/16021867713
[05:03:14] [PASSED] gt_was/14019449301
[05:03:14] [PASSED] gt_was/16028005424
[05:03:14] [PASSED] gt_was/14026578760
[05:03:14] [PASSED] gt_was/1409420604
[05:03:14] [PASSED] gt_was/1408615072
[05:03:14] [PASSED] gt_was/22010523718
[05:03:14] [PASSED] gt_was/14011006942
[05:03:14] [PASSED] gt_was/14014830051
[05:03:14] [PASSED] gt_was/18018781329
[05:03:14] [PASSED] gt_was/1509235366
[05:03:14] [PASSED] gt_was/18018781329
[05:03:14] [PASSED] gt_was/16016694945
[05:03:14] [PASSED] gt_was/14018575942
[05:03:14] [PASSED] gt_was/22016670082
[05:03:14] [PASSED] gt_was/22016670082
[05:03:14] [PASSED] gt_was/14017421178
[05:03:14] [PASSED] gt_was/16025250150
[05:03:14] [PASSED] gt_was/14021871409
[05:03:14] [PASSED] gt_was/16021865536
[05:03:14] [PASSED] gt_was/14021486841
[05:03:14] [PASSED] gt_was/14025160223
[05:03:14] [PASSED] gt_was/14026144927, 16029437861, 14026127056
[05:03:14] [PASSED] gt_was/14025635424
[05:03:14] [PASSED] gt_was/16028005424
[05:03:14] ============== [PASSED] xe_rtp_table_gt_test ===============
[05:03:14] ================== xe_rtp_table_gt_test ===================
[05:03:14] [PASSED] gt_tunings/Tuning: Blend Fill Caching Optimization Disable
[05:03:14] [PASSED] gt_tunings/Tuning: 32B Access Enable
[05:03:14] [PASSED] gt_tunings/Tuning: L3 cache
[05:03:14] [PASSED] gt_tunings/Tuning: L3 cache - media
[05:03:14] [PASSED] gt_tunings/Tuning: Compression Overfetch
[05:03:14] [PASSED] gt_tunings/Tuning: Compression Overfetch - media
[05:03:14] [PASSED] gt_tunings/Tuning: Enable compressible partial write overfetch in L3
[05:03:14] [PASSED] gt_tunings/Tuning: Enable compressible partial write overfetch in L3 - media
[05:03:14] [PASSED] gt_tunings/Tuning: L2 Overfetch Compressible Only
[05:03:14] [PASSED] gt_tunings/Tuning: L2 Overfetch Compressible Only - media
[05:03:14] [PASSED] gt_tunings/Tuning: Stateless compression control
[05:03:14] [PASSED] gt_tunings/Tuning: Stateless compression control - media
[05:03:14] [PASSED] gt_tunings/Tuning: L3 RW flush all Cache
[05:03:14] [PASSED] gt_tunings/Tuning: L3 RW flush all cache - media
[05:03:14] [PASSED] gt_tunings/Tuning: Set STLB Bank Hash Mode to 4KB
[05:03:14] ============== [PASSED] xe_rtp_table_gt_test ===============
[05:03:14] ================== xe_rtp_table_oob_test ==================
[05:03:14] [PASSED] oob_was/1607983814
[05:03:14] [PASSED] oob_was/16010904313
[05:03:14] [PASSED] oob_was/18022495364
[05:03:14] [PASSED] oob_was/22012773006
[05:03:14] [PASSED] oob_was/14014475959
[05:03:14] [PASSED] oob_was/22011391025
[05:03:14] [PASSED] oob_was/22012727170
[05:03:14] [PASSED] oob_was/22012727685
[05:03:14] [PASSED] oob_was/22016596838
[05:03:14] [PASSED] oob_was/18020744125
[05:03:14] [PASSED] oob_was/1409600907
[05:03:14] [PASSED] oob_was/22014953428
[05:03:14] [PASSED] oob_was/16017236439
[05:03:14] [PASSED] oob_was/14019821291
[05:03:14] [PASSED] oob_was/14015076503
[05:03:14] [PASSED] oob_was/22016122933
[05:03:14] [PASSED] oob_was/14018913170
[05:03:14] [PASSED] oob_was/14018094691
[05:03:14] [PASSED] oob_was/18024947630
[05:03:14] [PASSED] oob_was/16022287689
[05:03:14] [PASSED] oob_was/13011645652
[05:03:14] [PASSED] oob_was/14022293748
[05:03:14] [PASSED] oob_was/22019794406
[05:03:14] [PASSED] oob_was/22019338487
[05:03:14] [PASSED] oob_was/16023588340
[05:03:14] [PASSED] oob_was/14019789679
[05:03:14] [PASSED] oob_was/14022866841
[05:03:14] [PASSED] oob_was/16021333562
[05:03:14] [PASSED] oob_was/14016712196
[05:03:14] [PASSED] oob_was/14015568240
[05:03:14] [PASSED] oob_was/18013179988
[05:03:14] [PASSED] oob_was/1508761755
[05:03:14] [PASSED] oob_was/16023105232
[05:03:14] [PASSED] oob_was/16026508708
[05:03:14] [PASSED] oob_was/14020001231
[05:03:14] [PASSED] oob_was/16023683509
[05:03:14] [PASSED] oob_was/14025515070
[05:03:14] [PASSED] oob_was/15015404425_disable
[05:03:14] [PASSED] oob_was/16026007364
[05:03:14] [PASSED] oob_was/14020316580
[05:03:14] [PASSED] oob_was/14025883347
[05:03:14] [PASSED] oob_was/16029380221
[05:03:14] [PASSED] oob_was/22022079272
[05:03:14] [PASSED] oob_was/16029897822
[05:03:14] [PASSED] oob_was/14027054324
[05:03:14] [PASSED] oob_was/14025941587
[05:03:14] ============== [PASSED] xe_rtp_table_oob_test ==============
[05:03:14] ================ xe_rtp_table_dev_oob_test ================
[05:03:14] [PASSED] device_oob_was/22010954014
[05:03:14] [PASSED] device_oob_was/15015404425
[05:03:14] [PASSED] device_oob_was/22019338487_display
[05:03:14] [PASSED] device_oob_was/14022085890
[05:03:14] [PASSED] device_oob_was/14026539277
[05:03:14] [PASSED] device_oob_was/14026633728
[05:03:14] [PASSED] device_oob_was/14026746987
[05:03:14] [PASSED] device_oob_was/14026779378
[05:03:14] ============ [PASSED] xe_rtp_table_dev_oob_test ============
[05:03:14] ========== xe_rtp_table_missing_upper_bound_test ==========
[05:03:14] [PASSED] register_whitelist/WaAllowPMDepthAndInvocationCountAccessFromUMD, 1408556865
[05:03:14] [PASSED] register_whitelist/1508744258, 14012131227, 1808121037
[05:03:14] [PASSED] register_whitelist/1806527549
[05:03:14] [PASSED] register_whitelist/allow_read_ctx_timestamp
[05:03:14] [PASSED] register_whitelist/allow_read_queue_timestamp
[05:03:14] [PASSED] register_whitelist/16014440446
[05:03:14] [PASSED] register_whitelist/16017236439
[05:03:14] [PASSED] register_whitelist/16020183090
[05:03:14] [PASSED] register_whitelist/14024997852
[05:03:14] [PASSED] register_whitelist/14024997852
[05:03:14] ====== [PASSED] xe_rtp_table_missing_upper_bound_test ======
[05:03:14] =============== [PASSED] xe_rtp_tables_test ================
[05:03:14] =================== xe_rtp (3 subtests) ====================
[05:03:14] =================== xe_rtp_rules_tests ====================
[05:03:14] [PASSED] no
[05:03:14] [PASSED] yes
[05:03:14] [PASSED] no-and-no
[05:03:14] [PASSED] no-and-yes
[05:03:14] [PASSED] yes-and-no
[05:03:14] [PASSED] yes-and-yes
[05:03:14] [PASSED] no-or-no
[05:03:14] [PASSED] no-or-yes
[05:03:14] [PASSED] yes-or-no
[05:03:14] [PASSED] yes-or-yes
[05:03:14] [PASSED] no-yes-or-yes-no
[05:03:14] [PASSED] no-yes-or-yes-yes
[05:03:14] [PASSED] yes-yes-or-no-yes
[05:03:14] [PASSED] yes-yes-or-yes-yes
[05:03:14] [PASSED] no-no-or-yes-or-no
[05:03:14] [PASSED] or
[05:03:14] [PASSED] or-yes
[05:03:14] [PASSED] or-no
[05:03:14] [PASSED] yes-or
[05:03:14] [PASSED] no-or
[05:03:14] [PASSED] no-or-or-yes
[05:03:14] [PASSED] yes-or-or-no
[05:03:14] [PASSED] no-or-or-no
[05:03:14] [PASSED] missing-context-engine-class
[05:03:14] [PASSED] missing-context-engine-class-or-yes
[05:03:14] [PASSED] missing-context-engine-class-or-or-yes
[05:03:14] =============== [PASSED] xe_rtp_rules_tests ================
[05:03:14] =============== xe_rtp_process_to_sr_tests ================
[05:03:14] [PASSED] coalesce-same-reg
[05:03:14] [PASSED] coalesce-same-reg-literal-and-func
[05:03:14] [PASSED] no-match-no-add
[05:03:14] [PASSED] two-regs-two-entries
[05:03:14] [PASSED] clr-one-set-other
[05:03:14] [PASSED] set-field
[05:03:14] [PASSED] conflict-duplicate
[05:03:14] [PASSED] conflict-not-disjoint
[05:03:14] [PASSED] conflict-not-disjoint-literal-and-func
[05:03:14] [PASSED] conflict-reg-type
[05:03:14] [PASSED] bad-mcr-reg-forced-to-regular
[05:03:14] [PASSED] bad-regular-reg-forced-to-mcr
[05:03:14] =========== [PASSED] xe_rtp_process_to_sr_tests ============
[05:03:14] ================== xe_rtp_process_tests ===================
[05:03:14] [PASSED] active1
[05:03:14] [PASSED] active2
[05:03:14] [PASSED] active-inactive
[05:03:14] [PASSED] inactive-active
[05:03:14] [PASSED] inactive-active-inactive
[05:03:14] [PASSED] inactive-inactive-inactive
[05:03:14] ============== [PASSED] xe_rtp_process_tests ===============
[05:03:14] ===================== [PASSED] xe_rtp ======================
[05:03:14] ==================== xe_wa (1 subtest) =====================
[05:03:14] ======================== xe_wa_gt =========================
[05:03:14] [PASSED] TIGERLAKE B0
[05:03:14] [PASSED] DG1 A0
[05:03:14] [PASSED] DG1 B0
[05:03:14] [PASSED] ALDERLAKE_S A0
[05:03:14] [PASSED] ALDERLAKE_S B0
[05:03:14] [PASSED] ALDERLAKE_S C0
[05:03:14] [PASSED] ALDERLAKE_S D0
[05:03:14] [PASSED] ALDERLAKE_P A0
[05:03:14] [PASSED] ALDERLAKE_P B0
[05:03:14] [PASSED] ALDERLAKE_P C0
[05:03:14] [PASSED] ALDERLAKE_S RPLS D0
[05:03:14] [PASSED] ALDERLAKE_P RPLU E0
[05:03:14] [PASSED] DG2 G10 C0
[05:03:14] [PASSED] DG2 G11 B1
[05:03:14] [PASSED] DG2 G12 A1
[05:03:14] [PASSED] METEORLAKE 12.70(Xe_LPG) A0 13.00(Xe_LPM+) A0
[05:03:14] [PASSED] METEORLAKE 12.71(Xe_LPG) A0 13.00(Xe_LPM+) A0
[05:03:14] [PASSED] METEORLAKE 12.74(Xe_LPG+) A0 13.00(Xe_LPM+) A0
[05:03:14] [PASSED] LUNARLAKE 20.04(Xe2_LPG) A0 20.00(Xe2_LPM) A0
[05:03:14] [PASSED] LUNARLAKE 20.04(Xe2_LPG) B0 20.00(Xe2_LPM) A0
[05:03:14] [PASSED] BATTLEMAGE 20.01(Xe2_HPG) A0 13.01(Xe2_HPM) A1
[05:03:14] [PASSED] PANTHERLAKE 30.00(Xe3_LPG) A0 30.00(Xe3_LPM) A0
[05:03:14] ==================== [PASSED] xe_wa_gt =====================
[05:03:14] ====================== [PASSED] xe_wa ======================
[05:03:14] ============================================================
[05:03:14] Testing complete. Ran 821 tests: passed: 793, skipped: 28
[05:03:14] Elapsed time: 35.961s total, 1.802s configuring, 33.393s building, 0.711s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/tests/.kunitconfig
[05:03:14] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[05:03:16] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[05:03:41] Starting KUnit Kernel (1/1)...
[05:03:41] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[05:03:41] ============= refcount_interrupt (4 subtests) ==============
[05:03:41] [PASSED] test_single_irq_change
[05:03:41] [PASSED] test_nested_irq_change
[05:03:41] [PASSED] test_multiple_irq_change
[05:03:41] [PASSED] test_irq_save
[05:03:41] =============== [PASSED] refcount_interrupt ================
[05:03:41] ============ drm_test_pick_cmdline (2 subtests) ============
[05:03:41] [PASSED] drm_test_pick_cmdline_res_1920_1080_60
[05:03:41] =============== drm_test_pick_cmdline_named ===============
[05:03:41] [PASSED] NTSC
[05:03:41] [PASSED] NTSC-J
[05:03:41] [PASSED] PAL
[05:03:41] [PASSED] PAL-M
[05:03:41] =========== [PASSED] drm_test_pick_cmdline_named ===========
[05:03:41] ============== [PASSED] drm_test_pick_cmdline ==============
[05:03:41] == drm_test_atomic_get_connector_for_encoder (1 subtest) ===
[05:03:41] [PASSED] drm_test_drm_atomic_get_connector_for_encoder
[05:03:41] ==== [PASSED] drm_test_atomic_get_connector_for_encoder ====
[05:03:41] =========== drm_validate_clone_mode (2 subtests) ===========
[05:03:41] ============== drm_test_check_in_clone_mode ===============
[05:03:41] [PASSED] in_clone_mode
[05:03:41] [PASSED] not_in_clone_mode
[05:03:41] ========== [PASSED] drm_test_check_in_clone_mode ===========
[05:03:41] =============== drm_test_check_valid_clones ===============
[05:03:41] [PASSED] not_in_clone_mode
[05:03:41] [PASSED] valid_clone
[05:03:41] [PASSED] invalid_clone
[05:03:41] =========== [PASSED] drm_test_check_valid_clones ===========
[05:03:41] ============= [PASSED] drm_validate_clone_mode =============
[05:03:41] ============= drm_validate_modeset (1 subtest) =============
[05:03:41] [PASSED] drm_test_check_connector_changed_modeset
[05:03:41] ============== [PASSED] drm_validate_modeset ===============
[05:03:41] ====== drm_test_bridge_get_current_state (1 subtest) =======
[05:03:41] [PASSED] drm_test_drm_bridge_get_current_state_atomic
[05:03:41] ======== [PASSED] drm_test_bridge_get_current_state ========
[05:03:41] ====== drm_test_bridge_helper_reset_crtc (3 subtests) ======
[05:03:41] [PASSED] drm_test_drm_bridge_helper_reset_crtc_atomic
[05:03:41] [PASSED] drm_test_drm_bridge_helper_reset_crtc_atomic_disabled
[05:03:41] [PASSED] drm_test_drm_bridge_helper_hdmi_output_bus_fmts
[05:03:41] ======== [PASSED] drm_test_bridge_helper_reset_crtc ========
[05:03:41] ============== drm_bridge_alloc (2 subtests) ===============
[05:03:41] [PASSED] drm_test_drm_bridge_alloc_basic
[05:03:41] [PASSED] drm_test_drm_bridge_alloc_get_put
[05:03:41] ================ [PASSED] drm_bridge_alloc =================
[05:03:41] ============= drm_bridge_bus_fmt (5 subtests) ==============
[05:03:41] [PASSED] drm_test_bridge_rgb_yuv_rgb
[05:03:41] [PASSED] drm_test_bridge_must_convert_to_yuv444
[05:03:41] [PASSED] drm_test_bridge_hdmi_auto_rgb
[05:03:41] [PASSED] drm_test_bridge_auto_first
[05:03:41] [PASSED] drm_test_bridge_rgb_yuv_no_path
[05:03:41] =============== [PASSED] drm_bridge_bus_fmt ================
[05:03:41] ============= drm_cmdline_parser (40 subtests) =============
[05:03:41] [PASSED] drm_test_cmdline_force_d_only
[05:03:41] [PASSED] drm_test_cmdline_force_D_only_dvi
[05:03:41] [PASSED] drm_test_cmdline_force_D_only_hdmi
[05:03:41] [PASSED] drm_test_cmdline_force_D_only_not_digital
[05:03:41] [PASSED] drm_test_cmdline_force_e_only
[05:03:41] [PASSED] drm_test_cmdline_res
[05:03:41] [PASSED] drm_test_cmdline_res_vesa
[05:03:41] [PASSED] drm_test_cmdline_res_vesa_rblank
[05:03:41] [PASSED] drm_test_cmdline_res_rblank
[05:03:41] [PASSED] drm_test_cmdline_res_bpp
[05:03:41] [PASSED] drm_test_cmdline_res_refresh
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_interlaced
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_margins
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_force_off
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on_analog
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on_digital
[05:03:41] [PASSED] drm_test_cmdline_res_bpp_refresh_interlaced_margins_force_on
[05:03:41] [PASSED] drm_test_cmdline_res_margins_force_on
[05:03:41] [PASSED] drm_test_cmdline_res_vesa_margins
[05:03:41] [PASSED] drm_test_cmdline_name
[05:03:41] [PASSED] drm_test_cmdline_name_bpp
[05:03:41] [PASSED] drm_test_cmdline_name_option
[05:03:41] [PASSED] drm_test_cmdline_name_bpp_option
[05:03:41] [PASSED] drm_test_cmdline_rotate_0
[05:03:41] [PASSED] drm_test_cmdline_rotate_90
[05:03:41] [PASSED] drm_test_cmdline_rotate_180
[05:03:41] [PASSED] drm_test_cmdline_rotate_270
[05:03:41] [PASSED] drm_test_cmdline_hmirror
[05:03:41] [PASSED] drm_test_cmdline_vmirror
[05:03:41] [PASSED] drm_test_cmdline_margin_options
[05:03:41] [PASSED] drm_test_cmdline_multiple_options
[05:03:41] [PASSED] drm_test_cmdline_bpp_extra_and_option
[05:03:41] [PASSED] drm_test_cmdline_extra_and_option
[05:03:41] [PASSED] drm_test_cmdline_freestanding_options
[05:03:41] [PASSED] drm_test_cmdline_freestanding_force_e_and_options
[05:03:41] [PASSED] drm_test_cmdline_panel_orientation
[05:03:41] ================ drm_test_cmdline_invalid =================
[05:03:41] [PASSED] margin_only
[05:03:41] [PASSED] interlace_only
[05:03:41] [PASSED] res_missing_x
[05:03:41] [PASSED] res_missing_y
[05:03:41] [PASSED] res_bad_y
[05:03:41] [PASSED] res_missing_y_bpp
[05:03:41] [PASSED] res_bad_bpp
[05:03:41] [PASSED] res_bad_refresh
[05:03:41] [PASSED] res_bpp_refresh_force_on_off
[05:03:41] [PASSED] res_invalid_mode
[05:03:41] [PASSED] res_bpp_wrong_place_mode
[05:03:41] [PASSED] name_bpp_refresh
[05:03:41] [PASSED] name_refresh
[05:03:41] [PASSED] name_refresh_wrong_mode
[05:03:41] [PASSED] name_refresh_invalid_mode
[05:03:41] [PASSED] rotate_multiple
[05:03:41] [PASSED] rotate_invalid_val
[05:03:41] [PASSED] rotate_truncated
[05:03:41] [PASSED] invalid_option
[05:03:41] [PASSED] invalid_tv_option
[05:03:41] [PASSED] truncated_tv_option
[05:03:41] ============ [PASSED] drm_test_cmdline_invalid =============
[05:03:41] =============== drm_test_cmdline_tv_options ===============
[05:03:41] [PASSED] NTSC
[05:03:41] [PASSED] NTSC_443
[05:03:41] [PASSED] NTSC_J
[05:03:41] [PASSED] PAL
[05:03:41] [PASSED] PAL_M
[05:03:41] [PASSED] PAL_N
[05:03:41] [PASSED] SECAM
[05:03:41] [PASSED] MONO_525
[05:03:41] [PASSED] MONO_625
[05:03:41] =========== [PASSED] drm_test_cmdline_tv_options ===========
[05:03:41] =============== [PASSED] drm_cmdline_parser ================
[05:03:41] ========== drmm_connector_hdmi_init (20 subtests) ==========
[05:03:41] [PASSED] drm_test_connector_hdmi_init_valid
[05:03:41] [PASSED] drm_test_connector_hdmi_init_bpc_8
[05:03:41] [PASSED] drm_test_connector_hdmi_init_bpc_10
[05:03:41] [PASSED] drm_test_connector_hdmi_init_bpc_12
[05:03:41] [PASSED] drm_test_connector_hdmi_init_bpc_invalid
[05:03:41] [PASSED] drm_test_connector_hdmi_init_bpc_null
[05:03:41] [PASSED] drm_test_connector_hdmi_init_formats_empty
[05:03:41] [PASSED] drm_test_connector_hdmi_init_formats_no_rgb
[05:03:41] === drm_test_connector_hdmi_init_formats_yuv420_allowed ===
[05:03:41] [PASSED] supported_formats=0x9 yuv420_allowed=1
[05:03:41] [PASSED] supported_formats=0x9 yuv420_allowed=0
[05:03:41] [PASSED] supported_formats=0x5 yuv420_allowed=1
[05:03:41] [PASSED] supported_formats=0x5 yuv420_allowed=0
[05:03:41] === [PASSED] drm_test_connector_hdmi_init_formats_yuv420_allowed ===
[05:03:41] [PASSED] drm_test_connector_hdmi_init_null_ddc
[05:03:41] [PASSED] drm_test_connector_hdmi_init_null_product
[05:03:41] [PASSED] drm_test_connector_hdmi_init_null_vendor
[05:03:41] [PASSED] drm_test_connector_hdmi_init_product_length_exact
[05:03:41] [PASSED] drm_test_connector_hdmi_init_product_length_too_long
[05:03:41] [PASSED] drm_test_connector_hdmi_init_product_valid
[05:03:41] [PASSED] drm_test_connector_hdmi_init_vendor_length_exact
[05:03:41] [PASSED] drm_test_connector_hdmi_init_vendor_length_too_long
[05:03:41] [PASSED] drm_test_connector_hdmi_init_vendor_valid
[05:03:41] ========= drm_test_connector_hdmi_init_type_valid =========
[05:03:41] [PASSED] HDMI-A
[05:03:41] [PASSED] HDMI-B
[05:03:41] ===== [PASSED] drm_test_connector_hdmi_init_type_valid =====
[05:03:41] ======== drm_test_connector_hdmi_init_type_invalid ========
[05:03:41] [PASSED] Unknown
[05:03:41] [PASSED] VGA
[05:03:41] [PASSED] DVI-I
[05:03:41] [PASSED] DVI-D
[05:03:41] [PASSED] DVI-A
[05:03:41] [PASSED] Composite
[05:03:41] [PASSED] SVIDEO
[05:03:41] [PASSED] LVDS
[05:03:41] [PASSED] Component
[05:03:41] [PASSED] DIN
[05:03:41] [PASSED] DP
[05:03:41] [PASSED] TV
[05:03:41] [PASSED] eDP
[05:03:41] [PASSED] Virtual
[05:03:41] [PASSED] DSI
[05:03:41] [PASSED] DPI
[05:03:41] [PASSED] Writeback
[05:03:41] [PASSED] SPI
[05:03:41] [PASSED] USB
[05:03:41] ==== [PASSED] drm_test_connector_hdmi_init_type_invalid ====
[05:03:41] ============ [PASSED] drmm_connector_hdmi_init =============
[05:03:41] ============= drmm_connector_init (3 subtests) =============
[05:03:41] [PASSED] drm_test_drmm_connector_init
[05:03:41] [PASSED] drm_test_drmm_connector_init_null_ddc
[05:03:41] ========= drm_test_drmm_connector_init_type_valid =========
[05:03:41] [PASSED] Unknown
[05:03:41] [PASSED] VGA
[05:03:41] [PASSED] DVI-I
[05:03:41] [PASSED] DVI-D
[05:03:41] [PASSED] DVI-A
[05:03:41] [PASSED] Composite
[05:03:41] [PASSED] SVIDEO
[05:03:41] [PASSED] LVDS
[05:03:41] [PASSED] Component
[05:03:41] [PASSED] DIN
[05:03:41] [PASSED] DP
[05:03:41] [PASSED] HDMI-A
[05:03:41] [PASSED] HDMI-B
[05:03:41] [PASSED] TV
[05:03:41] [PASSED] eDP
[05:03:41] [PASSED] Virtual
[05:03:41] [PASSED] DSI
[05:03:41] [PASSED] DPI
[05:03:41] [PASSED] Writeback
[05:03:41] [PASSED] SPI
[05:03:41] [PASSED] USB
[05:03:41] ===== [PASSED] drm_test_drmm_connector_init_type_valid =====
[05:03:41] =============== [PASSED] drmm_connector_init ===============
[05:03:41] ========= drm_connector_dynamic_init (6 subtests) ==========
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_init
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_init_null_ddc
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_init_not_added
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_init_properties
[05:03:41] ===== drm_test_drm_connector_dynamic_init_type_valid ======
[05:03:41] [PASSED] Unknown
[05:03:41] [PASSED] VGA
[05:03:41] [PASSED] DVI-I
[05:03:41] [PASSED] DVI-D
[05:03:41] [PASSED] DVI-A
[05:03:41] [PASSED] Composite
[05:03:41] [PASSED] SVIDEO
[05:03:41] [PASSED] LVDS
[05:03:41] [PASSED] Component
[05:03:41] [PASSED] DIN
[05:03:41] [PASSED] DP
[05:03:41] [PASSED] HDMI-A
[05:03:41] [PASSED] HDMI-B
[05:03:41] [PASSED] TV
[05:03:41] [PASSED] eDP
[05:03:41] [PASSED] Virtual
[05:03:41] [PASSED] DSI
[05:03:41] [PASSED] DPI
[05:03:41] [PASSED] Writeback
[05:03:41] [PASSED] SPI
[05:03:41] [PASSED] USB
[05:03:41] = [PASSED] drm_test_drm_connector_dynamic_init_type_valid ==
[05:03:41] ======== drm_test_drm_connector_dynamic_init_name =========
[05:03:41] [PASSED] Unknown
[05:03:41] [PASSED] VGA
[05:03:41] [PASSED] DVI-I
[05:03:41] [PASSED] DVI-D
[05:03:41] [PASSED] DVI-A
[05:03:41] [PASSED] Composite
[05:03:41] [PASSED] SVIDEO
[05:03:41] [PASSED] LVDS
[05:03:41] [PASSED] Component
[05:03:41] [PASSED] DIN
[05:03:41] [PASSED] DP
[05:03:41] [PASSED] HDMI-A
[05:03:41] [PASSED] HDMI-B
[05:03:41] [PASSED] TV
[05:03:41] [PASSED] eDP
[05:03:41] [PASSED] Virtual
[05:03:41] [PASSED] DSI
[05:03:41] [PASSED] DPI
[05:03:41] [PASSED] Writeback
[05:03:41] [PASSED] SPI
[05:03:41] [PASSED] USB
[05:03:41] ==== [PASSED] drm_test_drm_connector_dynamic_init_name =====
[05:03:41] =========== [PASSED] drm_connector_dynamic_init ============
[05:03:41] ==== drm_connector_dynamic_register_early (4 subtests) =====
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_early_on_list
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_early_defer
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_early_no_init
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_early_no_mode_object
[05:03:41] ====== [PASSED] drm_connector_dynamic_register_early =======
[05:03:41] ======= drm_connector_dynamic_register (7 subtests) ========
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_on_list
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_no_defer
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_no_init
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_mode_object
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_sysfs
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_sysfs_name
[05:03:41] [PASSED] drm_test_drm_connector_dynamic_register_debugfs
[05:03:41] ========= [PASSED] drm_connector_dynamic_register ==========
[05:03:41] = drm_connector_attach_broadcast_rgb_property (2 subtests) =
[05:03:41] [PASSED] drm_test_drm_connector_attach_broadcast_rgb_property
[05:03:41] [PASSED] drm_test_drm_connector_attach_broadcast_rgb_property_hdmi_connector
[05:03:41] === [PASSED] drm_connector_attach_broadcast_rgb_property ===
[05:03:41] ========== drm_get_tv_mode_from_name (2 subtests) ==========
[05:03:41] ========== drm_test_get_tv_mode_from_name_valid ===========
[05:03:41] [PASSED] NTSC
[05:03:41] [PASSED] NTSC-443
[05:03:41] [PASSED] NTSC-J
[05:03:41] [PASSED] PAL
[05:03:41] [PASSED] PAL-M
[05:03:41] [PASSED] PAL-N
[05:03:41] [PASSED] SECAM
[05:03:41] [PASSED] Mono
[05:03:41] ====== [PASSED] drm_test_get_tv_mode_from_name_valid =======
[05:03:41] [PASSED] drm_test_get_tv_mode_from_name_truncated
[05:03:41] ============ [PASSED] drm_get_tv_mode_from_name ============
[05:03:41] = drm_test_connector_hdmi_compute_mode_clock (12 subtests) =
[05:03:41] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb
[05:03:41] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_10bpc
[05:03:41] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_10bpc_vic_1
[05:03:41] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_12bpc
[05:03:41] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_12bpc_vic_1
[05:03:41] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_double
[05:03:41] = drm_test_connector_hdmi_compute_mode_clock_yuv420_valid =
[05:03:41] [PASSED] VIC 96
[05:03:41] [PASSED] VIC 97
[05:03:41] [PASSED] VIC 101
[05:03:41] [PASSED] VIC 102
[05:03:41] [PASSED] VIC 106
[05:03:41] [PASSED] VIC 107
[05:03:41] === [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_valid ===
[05:03:41] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_10_bpc
[05:03:41] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_12_bpc
[05:03:41] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_8_bpc
[05:03:41] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_10_bpc
[05:03:41] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_12_bpc
[05:03:41] === [PASSED] drm_test_connector_hdmi_compute_mode_clock ====
[05:03:41] == drm_hdmi_connector_get_broadcast_rgb_name (2 subtests) ==
[05:03:41] === drm_test_drm_hdmi_connector_get_broadcast_rgb_name ====
[05:03:41] [PASSED] Automatic
[05:03:41] [PASSED] Full
[05:03:41] [PASSED] Limited 16:235
[05:03:41] === [PASSED] drm_test_drm_hdmi_connector_get_broadcast_rgb_name ===
[05:03:41] [PASSED] drm_test_drm_hdmi_connector_get_broadcast_rgb_name_invalid
[05:03:41] ==== [PASSED] drm_hdmi_connector_get_broadcast_rgb_name ====
[05:03:41] == drm_hdmi_connector_get_output_format_name (2 subtests) ==
[05:03:41] === drm_test_drm_hdmi_connector_get_output_format_name ====
[05:03:41] [PASSED] RGB
[05:03:41] [PASSED] YUV 4:2:0
[05:03:41] [PASSED] YUV 4:2:2
[05:03:41] [PASSED] YUV 4:4:4
[05:03:41] === [PASSED] drm_test_drm_hdmi_connector_get_output_format_name ===
[05:03:41] [PASSED] drm_test_drm_hdmi_connector_get_output_format_name_invalid
[05:03:41] ==== [PASSED] drm_hdmi_connector_get_output_format_name ====
[05:03:41] ============= drm_damage_helper (21 subtests) ==============
[05:03:41] [PASSED] drm_test_damage_iter_no_damage
[05:03:41] [PASSED] drm_test_damage_iter_no_damage_fractional_src
[05:03:41] [PASSED] drm_test_damage_iter_no_damage_src_moved
[05:03:41] [PASSED] drm_test_damage_iter_no_damage_fractional_src_moved
[05:03:41] [PASSED] drm_test_damage_iter_no_damage_not_visible
[05:03:41] [PASSED] drm_test_damage_iter_no_damage_no_crtc
[05:03:41] [PASSED] drm_test_damage_iter_no_damage_no_fb
[05:03:41] [PASSED] drm_test_damage_iter_simple_damage
[05:03:41] [PASSED] drm_test_damage_iter_single_damage
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_intersect_src
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_outside_src
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_fractional_src
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_intersect_fractional_src
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_outside_fractional_src
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_src_moved
[05:03:41] [PASSED] drm_test_damage_iter_single_damage_fractional_src_moved
[05:03:41] [PASSED] drm_test_damage_iter_damage
[05:03:41] [PASSED] drm_test_damage_iter_damage_one_intersect
[05:03:41] [PASSED] drm_test_damage_iter_damage_one_outside
[05:03:41] [PASSED] drm_test_damage_iter_damage_src_moved
[05:03:41] [PASSED] drm_test_damage_iter_damage_not_visible
[05:03:41] ================ [PASSED] drm_damage_helper ================
[05:03:41] ============== drm_dp_mst_helper (3 subtests) ==============
[05:03:41] ============== drm_test_dp_mst_calc_pbn_mode ==============
[05:03:41] [PASSED] Clock 154000 BPP 30 DSC disabled
[05:03:41] [PASSED] Clock 234000 BPP 30 DSC disabled
[05:03:41] [PASSED] Clock 297000 BPP 24 DSC disabled
[05:03:41] [PASSED] Clock 332880 BPP 24 DSC enabled
[05:03:41] [PASSED] Clock 324540 BPP 24 DSC enabled
[05:03:41] ========== [PASSED] drm_test_dp_mst_calc_pbn_mode ==========
[05:03:41] ============== drm_test_dp_mst_calc_pbn_div ===============
[05:03:41] [PASSED] Link rate 2000000 lane count 4
[05:03:41] [PASSED] Link rate 2000000 lane count 2
[05:03:41] [PASSED] Link rate 2000000 lane count 1
[05:03:41] [PASSED] Link rate 1350000 lane count 4
[05:03:41] [PASSED] Link rate 1350000 lane count 2
[05:03:41] [PASSED] Link rate 1350000 lane count 1
[05:03:41] [PASSED] Link rate 1000000 lane count 4
[05:03:41] [PASSED] Link rate 1000000 lane count 2
[05:03:41] [PASSED] Link rate 1000000 lane count 1
[05:03:41] [PASSED] Link rate 810000 lane count 4
[05:03:41] [PASSED] Link rate 810000 lane count 2
[05:03:41] [PASSED] Link rate 810000 lane count 1
[05:03:41] [PASSED] Link rate 540000 lane count 4
[05:03:41] [PASSED] Link rate 540000 lane count 2
[05:03:41] [PASSED] Link rate 540000 lane count 1
[05:03:41] [PASSED] Link rate 270000 lane count 4
[05:03:41] [PASSED] Link rate 270000 lane count 2
[05:03:41] [PASSED] Link rate 270000 lane count 1
[05:03:41] [PASSED] Link rate 162000 lane count 4
[05:03:41] [PASSED] Link rate 162000 lane count 2
[05:03:41] [PASSED] Link rate 162000 lane count 1
[05:03:41] ========== [PASSED] drm_test_dp_mst_calc_pbn_div ===========
[05:03:41] ========= drm_test_dp_mst_sideband_msg_req_decode =========
[05:03:41] [PASSED] DP_ENUM_PATH_RESOURCES with port number
[05:03:41] [PASSED] DP_POWER_UP_PHY with port number
[05:03:41] [PASSED] DP_POWER_DOWN_PHY with port number
[05:03:41] [PASSED] DP_ALLOCATE_PAYLOAD with SDP stream sinks
[05:03:41] [PASSED] DP_ALLOCATE_PAYLOAD with port number
[05:03:41] [PASSED] DP_ALLOCATE_PAYLOAD with VCPI
[05:03:41] [PASSED] DP_ALLOCATE_PAYLOAD with PBN
[05:03:41] [PASSED] DP_QUERY_PAYLOAD with port number
[05:03:41] [PASSED] DP_QUERY_PAYLOAD with VCPI
[05:03:41] [PASSED] DP_REMOTE_DPCD_READ with port number
[05:03:41] [PASSED] DP_REMOTE_DPCD_READ with DPCD address
[05:03:41] [PASSED] DP_REMOTE_DPCD_READ with max number of bytes
[05:03:41] [PASSED] DP_REMOTE_DPCD_WRITE with port number
[05:03:41] [PASSED] DP_REMOTE_DPCD_WRITE with DPCD address
[05:03:41] [PASSED] DP_REMOTE_DPCD_WRITE with data array
[05:03:41] [PASSED] DP_REMOTE_I2C_READ with port number
[05:03:41] [PASSED] DP_REMOTE_I2C_READ with I2C device ID
[05:03:41] [PASSED] DP_REMOTE_I2C_READ with transactions array
[05:03:41] [PASSED] DP_REMOTE_I2C_WRITE with port number
[05:03:41] [PASSED] DP_REMOTE_I2C_WRITE with I2C device ID
[05:03:41] [PASSED] DP_REMOTE_I2C_WRITE with data array
[05:03:41] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream ID
[05:03:41] [PASSED] DP_QUERY_STREAM_ENC_STATUS with client ID
[05:03:41] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream event
[05:03:41] [PASSED] DP_QUERY_STREAM_ENC_STATUS with valid stream event
[05:03:41] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream behavior
[05:03:41] [PASSED] DP_QUERY_STREAM_ENC_STATUS with a valid stream behavior
[05:03:41] ===== [PASSED] drm_test_dp_mst_sideband_msg_req_decode =====
[05:03:41] ================ [PASSED] drm_dp_mst_helper ================
[05:03:41] ================== drm_exec (7 subtests) ===================
[05:03:41] [PASSED] sanitycheck
[05:03:41] [PASSED] test_lock
[05:03:41] [PASSED] test_lock_unlock
[05:03:41] [PASSED] test_duplicates
[05:03:41] [PASSED] test_prepare
[05:03:41] [PASSED] test_prepare_array
[05:03:41] [PASSED] test_multiple_loops
[05:03:41] ==================== [PASSED] drm_exec =====================
[05:03:41] =========== drm_format_helper_test (17 subtests) ===========
[05:03:41] ============== drm_test_fb_xrgb8888_to_gray8 ==============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ========== [PASSED] drm_test_fb_xrgb8888_to_gray8 ==========
[05:03:41] ============= drm_test_fb_xrgb8888_to_rgb332 ==============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb332 ==========
[05:03:41] ============= drm_test_fb_xrgb8888_to_rgb565 ==============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb565 ==========
[05:03:41] ============ drm_test_fb_xrgb8888_to_xrgb1555 =============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======== [PASSED] drm_test_fb_xrgb8888_to_xrgb1555 =========
[05:03:41] ============ drm_test_fb_xrgb8888_to_argb1555 =============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======== [PASSED] drm_test_fb_xrgb8888_to_argb1555 =========
[05:03:41] ============ drm_test_fb_xrgb8888_to_rgba5551 =============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======== [PASSED] drm_test_fb_xrgb8888_to_rgba5551 =========
[05:03:41] ============= drm_test_fb_xrgb8888_to_rgb888 ==============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb888 ==========
[05:03:41] ============= drm_test_fb_xrgb8888_to_bgr888 ==============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ========= [PASSED] drm_test_fb_xrgb8888_to_bgr888 ==========
[05:03:41] ============ drm_test_fb_xrgb8888_to_argb8888 =============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======== [PASSED] drm_test_fb_xrgb8888_to_argb8888 =========
[05:03:41] =========== drm_test_fb_xrgb8888_to_xrgb2101010 ===========
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======= [PASSED] drm_test_fb_xrgb8888_to_xrgb2101010 =======
[05:03:41] =========== drm_test_fb_xrgb8888_to_argb2101010 ===========
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======= [PASSED] drm_test_fb_xrgb8888_to_argb2101010 =======
[05:03:41] ============== drm_test_fb_xrgb8888_to_mono ===============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ========== [PASSED] drm_test_fb_xrgb8888_to_mono ===========
[05:03:41] ==================== drm_test_fb_swab =====================
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ================ [PASSED] drm_test_fb_swab =================
[05:03:41] ============ drm_test_fb_xrgb8888_to_xbgr8888 =============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======== [PASSED] drm_test_fb_xrgb8888_to_xbgr8888 =========
[05:03:41] ============ drm_test_fb_xrgb8888_to_abgr8888 =============
[05:03:41] [PASSED] single_pixel_source_buffer
[05:03:41] [PASSED] single_pixel_clip_rectangle
[05:03:41] [PASSED] well_known_colors
[05:03:41] [PASSED] destination_pitch
[05:03:41] ======== [PASSED] drm_test_fb_xrgb8888_to_abgr8888 =========
[05:03:41] ================= drm_test_fb_clip_offset =================
[05:03:41] [PASSED] pass through
[05:03:41] [PASSED] horizontal offset
[05:03:41] [PASSED] vertical offset
[05:03:41] [PASSED] horizontal and vertical offset
[05:03:41] [PASSED] horizontal offset (custom pitch)
[05:03:41] [PASSED] vertical offset (custom pitch)
[05:03:41] [PASSED] horizontal and vertical offset (custom pitch)
[05:03:41] ============= [PASSED] drm_test_fb_clip_offset =============
[05:03:41] =================== drm_test_fb_memcpy ====================
[05:03:41] [PASSED] single_pixel_source_buffer: XR24 little-endian (0x34325258)
[05:03:41] [PASSED] single_pixel_source_buffer: XRA8 little-endian (0x38415258)
[05:03:41] [PASSED] single_pixel_source_buffer: YU24 little-endian (0x34325559)
[05:03:41] [PASSED] single_pixel_clip_rectangle: XB24 little-endian (0x34324258)
[05:03:41] [PASSED] single_pixel_clip_rectangle: XRA8 little-endian (0x38415258)
[05:03:41] [PASSED] single_pixel_clip_rectangle: YU24 little-endian (0x34325559)
[05:03:41] [PASSED] well_known_colors: XB24 little-endian (0x34324258)
[05:03:41] [PASSED] well_known_colors: XRA8 little-endian (0x38415258)
[05:03:41] [PASSED] well_known_colors: YU24 little-endian (0x34325559)
[05:03:41] [PASSED] destination_pitch: XB24 little-endian (0x34324258)
[05:03:41] [PASSED] destination_pitch: XRA8 little-endian (0x38415258)
[05:03:41] [PASSED] destination_pitch: YU24 little-endian (0x34325559)
[05:03:41] =============== [PASSED] drm_test_fb_memcpy ================
[05:03:41] ============= [PASSED] drm_format_helper_test ==============
[05:03:41] ================= drm_format (18 subtests) =================
[05:03:41] [PASSED] drm_test_format_block_width_invalid
[05:03:41] [PASSED] drm_test_format_block_width_one_plane
[05:03:41] [PASSED] drm_test_format_block_width_two_plane
[05:03:41] [PASSED] drm_test_format_block_width_three_plane
[05:03:41] [PASSED] drm_test_format_block_width_tiled
[05:03:41] [PASSED] drm_test_format_block_height_invalid
[05:03:41] [PASSED] drm_test_format_block_height_one_plane
[05:03:41] [PASSED] drm_test_format_block_height_two_plane
[05:03:41] [PASSED] drm_test_format_block_height_three_plane
[05:03:41] [PASSED] drm_test_format_block_height_tiled
[05:03:41] [PASSED] drm_test_format_min_pitch_invalid
[05:03:41] [PASSED] drm_test_format_min_pitch_one_plane_8bpp
[05:03:41] [PASSED] drm_test_format_min_pitch_one_plane_16bpp
[05:03:41] [PASSED] drm_test_format_min_pitch_one_plane_24bpp
[05:03:41] [PASSED] drm_test_format_min_pitch_one_plane_32bpp
[05:03:41] [PASSED] drm_test_format_min_pitch_two_plane
[05:03:41] [PASSED] drm_test_format_min_pitch_three_plane_8bpp
[05:03:41] [PASSED] drm_test_format_min_pitch_tiled
[05:03:41] =================== [PASSED] drm_format ====================
[05:03:41] ============== drm_framebuffer (10 subtests) ===============
[05:03:41] ========== drm_test_framebuffer_check_src_coords ==========
[05:03:41] [PASSED] Success: source fits into fb
[05:03:41] [PASSED] Fail: overflowing fb with x-axis coordinate
[05:03:41] [PASSED] Fail: overflowing fb with y-axis coordinate
[05:03:41] [PASSED] Fail: overflowing fb with source width
[05:03:41] [PASSED] Fail: overflowing fb with source height
[05:03:41] ====== [PASSED] drm_test_framebuffer_check_src_coords ======
[05:03:41] [PASSED] drm_test_framebuffer_cleanup
[05:03:41] =============== drm_test_framebuffer_create ===============
[05:03:41] [PASSED] ABGR8888 normal sizes
[05:03:41] [PASSED] ABGR8888 max sizes
[05:03:41] [PASSED] ABGR8888 pitch greater than min required
[05:03:41] [PASSED] ABGR8888 pitch less than min required
[05:03:41] [PASSED] ABGR8888 Invalid width
[05:03:41] [PASSED] ABGR8888 Invalid buffer handle
[05:03:41] [PASSED] No pixel format
[05:03:41] [PASSED] ABGR8888 Width 0
[05:03:41] [PASSED] ABGR8888 Height 0
[05:03:41] [PASSED] ABGR8888 Out of bound height * pitch combination
[05:03:41] [PASSED] ABGR8888 Large buffer offset
[05:03:41] [PASSED] ABGR8888 Buffer offset for inexistent plane
[05:03:41] [PASSED] ABGR8888 Invalid flag
[05:03:41] [PASSED] ABGR8888 Set DRM_MODE_FB_MODIFIERS without modifiers
[05:03:41] [PASSED] ABGR8888 Valid buffer modifier
[05:03:41] [PASSED] ABGR8888 Invalid buffer modifier(DRM_FORMAT_MOD_SAMSUNG_64_32_TILE)
[05:03:41] [PASSED] ABGR8888 Extra pitches without DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] ABGR8888 Extra pitches with DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] NV12 Normal sizes
[05:03:41] [PASSED] NV12 Max sizes
[05:03:41] [PASSED] NV12 Invalid pitch
[05:03:41] [PASSED] NV12 Invalid modifier/missing DRM_MODE_FB_MODIFIERS flag
[05:03:41] [PASSED] NV12 different modifier per-plane
[05:03:41] [PASSED] NV12 with DRM_FORMAT_MOD_SAMSUNG_64_32_TILE
[05:03:41] [PASSED] NV12 Valid modifiers without DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] NV12 Modifier for inexistent plane
[05:03:41] [PASSED] NV12 Handle for inexistent plane
[05:03:41] [PASSED] NV12 Handle for inexistent plane without DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] YVU420 DRM_MODE_FB_MODIFIERS set without modifier
[05:03:41] [PASSED] YVU420 Normal sizes
[05:03:41] [PASSED] YVU420 Max sizes
[05:03:41] [PASSED] YVU420 Invalid pitch
[05:03:41] [PASSED] YVU420 Different pitches
[05:03:41] [PASSED] YVU420 Different buffer offsets/pitches
[05:03:41] [PASSED] YVU420 Modifier set just for plane 0, without DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] YVU420 Modifier set just for planes 0, 1, without DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] YVU420 Modifier set just for plane 0, 1, with DRM_MODE_FB_MODIFIERS
[05:03:41] [PASSED] YVU420 Valid modifier
[05:03:41] [PASSED] YVU420 Different modifiers per plane
[05:03:41] [PASSED] YVU420 Modifier for inexistent plane
[05:03:41] [PASSED] YUV420_10BIT Invalid modifier(DRM_FORMAT_MOD_LINEAR)
[05:03:41] [PASSED] X0L2 Normal sizes
[05:03:41] [PASSED] X0L2 Max sizes
[05:03:41] [PASSED] X0L2 Invalid pitch
[05:03:41] [PASSED] X0L2 Pitch greater than minimum required
[05:03:41] [PASSED] X0L2 Handle for inexistent plane
[05:03:41] [PASSED] X0L2 Offset for inexistent plane, without DRM_MODE_FB_MODIFIERS set
[05:03:41] [PASSED] X0L2 Modifier without DRM_MODE_FB_MODIFIERS set
[05:03:41] [PASSED] X0L2 Valid modifier
[05:03:41] [PASSED] X0L2 Modifier for inexistent plane
[05:03:41] =========== [PASSED] drm_test_framebuffer_create ===========
[05:03:41] [PASSED] drm_test_framebuffer_free
[05:03:41] [PASSED] drm_test_framebuffer_init
[05:03:41] [PASSED] drm_test_framebuffer_init_bad_format
[05:03:41] [PASSED] drm_test_framebuffer_init_dev_mismatch
[05:03:41] [PASSED] drm_test_framebuffer_lookup
[05:03:41] [PASSED] drm_test_framebuffer_lookup_inexistent
[05:03:41] [PASSED] drm_test_framebuffer_modifiers_not_supported
[05:03:41] ================= [PASSED] drm_framebuffer =================
[05:03:41] ================ drm_gem_shmem (8 subtests) ================
[05:03:41] [PASSED] drm_gem_shmem_test_obj_create
[05:03:41] [PASSED] drm_gem_shmem_test_obj_create_private
[05:03:41] [PASSED] drm_gem_shmem_test_pin_pages
[05:03:41] [PASSED] drm_gem_shmem_test_vmap
[05:03:41] [PASSED] drm_gem_shmem_test_get_sg_table
[05:03:41] [PASSED] drm_gem_shmem_test_get_pages_sgt
[05:03:41] [PASSED] drm_gem_shmem_test_madvise
[05:03:41] [PASSED] drm_gem_shmem_test_purge
[05:03:41] ================== [PASSED] drm_gem_shmem ==================
[05:03:41] === drm_atomic_helper_connector_hdmi_check (29 subtests) ===
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_auto_cea_mode
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_auto_cea_mode_vic_1
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_full_cea_mode
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_full_cea_mode_vic_1
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_limited_cea_mode
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_limited_cea_mode_vic_1
[05:03:41] ====== drm_test_check_broadcast_rgb_cea_mode_yuv420 =======
[05:03:41] [PASSED] Automatic
[05:03:41] [PASSED] Full
[05:03:41] [PASSED] Limited 16:235
[05:03:41] == [PASSED] drm_test_check_broadcast_rgb_cea_mode_yuv420 ===
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_crtc_mode_changed
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_crtc_mode_not_changed
[05:03:41] [PASSED] drm_test_check_disable_connector
[05:03:41] [PASSED] drm_test_check_hdmi_funcs_reject_rate
[05:03:41] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_rgb
[05:03:41] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_yuv420
[05:03:41] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_ignore_yuv422
[05:03:41] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_ignore_yuv420
[05:03:41] [PASSED] drm_test_check_driver_unsupported_fallback_yuv420
[05:03:41] [PASSED] drm_test_check_output_bpc_crtc_mode_changed
[05:03:41] [PASSED] drm_test_check_output_bpc_crtc_mode_not_changed
[05:03:41] [PASSED] drm_test_check_output_bpc_dvi
[05:03:41] [PASSED] drm_test_check_output_bpc_format_vic_1
[05:03:41] [PASSED] drm_test_check_output_bpc_format_display_8bpc_only
[05:03:41] [PASSED] drm_test_check_output_bpc_format_display_rgb_only
[05:03:41] [PASSED] drm_test_check_output_bpc_format_driver_8bpc_only
[05:03:41] [PASSED] drm_test_check_output_bpc_format_driver_rgb_only
[05:03:41] [PASSED] drm_test_check_tmds_char_rate_rgb_8bpc
[05:03:41] [PASSED] drm_test_check_tmds_char_rate_rgb_10bpc
[05:03:41] [PASSED] drm_test_check_tmds_char_rate_rgb_12bpc
[05:03:41] ============ drm_test_check_hdmi_color_format =============
[05:03:41] [PASSED] AUTO -> RGB
[05:03:41] [PASSED] YCBCR422 -> YUV422
[05:03:41] [PASSED] YCBCR420 -> YUV420
[05:03:41] [PASSED] YCBCR444 -> YUV444
[05:03:41] [PASSED] RGB -> RGB
[05:03:41] ======== [PASSED] drm_test_check_hdmi_color_format =========
[05:03:41] ======== drm_test_check_hdmi_color_format_420_only ========
[05:03:41] [PASSED] RGB should fail
[05:03:41] [PASSED] YUV444 should fail
[05:03:41] [PASSED] YUV422 should fail
[05:03:41] [PASSED] YUV420 should work
[05:03:41] ==== [PASSED] drm_test_check_hdmi_color_format_420_only ====
[05:03:41] ===== [PASSED] drm_atomic_helper_connector_hdmi_check ======
[05:03:41] === drm_atomic_helper_connector_hdmi_reset (6 subtests) ====
[05:03:41] [PASSED] drm_test_check_broadcast_rgb_value
[05:03:41] [PASSED] drm_test_check_bpc_8_value
[05:03:41] [PASSED] drm_test_check_bpc_10_value
[05:03:41] [PASSED] drm_test_check_bpc_12_value
[05:03:41] [PASSED] drm_test_check_format_value
[05:03:41] [PASSED] drm_test_check_tmds_char_value
[05:03:41] ===== [PASSED] drm_atomic_helper_connector_hdmi_reset ======
[05:03:41] = drm_atomic_helper_connector_hdmi_mode_valid (7 subtests) =
[05:03:41] [PASSED] drm_test_check_mode_valid
[05:03:41] [PASSED] drm_test_check_mode_valid_reject
[05:03:41] [PASSED] drm_test_check_mode_valid_reject_rate
[05:03:41] [PASSED] drm_test_check_mode_valid_reject_max_clock
[05:03:41] [PASSED] drm_test_check_mode_valid_yuv420_only_max_clock
[05:03:41] [PASSED] drm_test_check_mode_valid_reject_yuv420_only_connector
[05:03:41] [PASSED] drm_test_check_mode_valid_accept_yuv420_also_connector_rgb
[05:03:41] === [PASSED] drm_atomic_helper_connector_hdmi_mode_valid ===
[05:03:41] = drm_atomic_helper_connector_hdmi_infoframes (5 subtests) =
[05:03:41] [PASSED] drm_test_check_infoframes
[05:03:41] [PASSED] drm_test_check_reject_avi_infoframe
[05:03:41] [PASSED] drm_test_check_reject_hdr_infoframe_bpc_8
[05:03:41] [PASSED] drm_test_check_reject_hdr_infoframe_bpc_10
[05:03:41] [PASSED] drm_test_check_reject_audio_infoframe
[05:03:41] === [PASSED] drm_atomic_helper_connector_hdmi_infoframes ===
[05:03:41] ================= drm_managed (2 subtests) =================
[05:03:41] [PASSED] drm_test_managed_release_action
[05:03:41] [PASSED] drm_test_managed_run_action
[05:03:41] =================== [PASSED] drm_managed ===================
[05:03:41] =================== drm_mm (6 subtests) ====================
[05:03:41] [PASSED] drm_test_mm_init
[05:03:41] [PASSED] drm_test_mm_debug
[05:03:41] [PASSED] drm_test_mm_align32
[05:03:41] [PASSED] drm_test_mm_align64
[05:03:41] [PASSED] drm_test_mm_lowest
[05:03:41] [PASSED] drm_test_mm_highest
[05:03:41] ===================== [PASSED] drm_mm ======================
[05:03:41] ============= drm_modes_analog_tv (5 subtests) =============
[05:03:41] [PASSED] drm_test_modes_analog_tv_mono_576i
[05:03:41] [PASSED] drm_test_modes_analog_tv_ntsc_480i
[05:03:41] [PASSED] drm_test_modes_analog_tv_ntsc_480i_inlined
[05:03:41] [PASSED] drm_test_modes_analog_tv_pal_576i
[05:03:41] [PASSED] drm_test_modes_analog_tv_pal_576i_inlined
[05:03:41] =============== [PASSED] drm_modes_analog_tv ===============
[05:03:41] ============== drm_panic_helper (3 subtests) ===============
[05:03:41] ============= drm_test_panic_screen_user_map ==============
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 494 x 494 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 RG16 little-endian (0x36314752)
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 RG24 little-endian (0x34324752)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 494 x 494 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG16 little-endian (0x36314752)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG24 little-endian (0x34324752)
[05:03:41] ========= [PASSED] drm_test_panic_screen_user_map ==========
[05:03:41] ============= drm_test_panic_screen_user_page =============
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 494 x 494 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 RG16 little-endian (0x36314752)
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 RG24 little-endian (0x34324752)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 494 x 494 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG16 little-endian (0x36314752)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG24 little-endian (0x34324752)
[05:03:41] ========= [PASSED] drm_test_panic_screen_user_page =========
[05:03:41] ========== drm_test_panic_screen_user_set_pixel ===========
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 494 x 494 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 RG16 little-endian (0x36314752)
[05:03:41] [PASSED] Panic screen user, mode: 1024 x 768 RG24 little-endian (0x34324752)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 494 x 494 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1920 x 1080 XR24 little-endian (0x34325258)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG16 little-endian (0x36314752)
[05:03:41] [PASSED] Panic screen kmsg, mode: 1024 x 768 RG24 little-endian (0x34324752)
[05:03:41] ====== [PASSED] drm_test_panic_screen_user_set_pixel =======
[05:03:41] ================ [PASSED] drm_panic_helper =================
[05:03:41] ============== drm_plane_helper (2 subtests) ===============
[05:03:41] =============== drm_test_check_plane_state ================
[05:03:41] [PASSED] clipping_simple
[05:03:41] [PASSED] clipping_rotate_reflect
[05:03:41] [PASSED] positioning_simple
[05:03:41] [PASSED] upscaling
[05:03:41] [PASSED] downscaling
[05:03:41] [PASSED] rounding1
[05:03:41] [PASSED] rounding2
[05:03:41] [PASSED] rounding3
[05:03:41] [PASSED] rounding4
[05:03:41] =========== [PASSED] drm_test_check_plane_state ============
[05:03:41] =========== drm_test_check_invalid_plane_state ============
[05:03:41] [PASSED] positioning_invalid
[05:03:41] [PASSED] upscaling_invalid
[05:03:41] [PASSED] downscaling_invalid
[05:03:41] ======= [PASSED] drm_test_check_invalid_plane_state ========
[05:03:41] ================ [PASSED] drm_plane_helper =================
[05:03:41] ====== drm_connector_helper_tv_get_modes (1 subtest) =======
[05:03:41] ====== drm_test_connector_helper_tv_get_modes_check =======
[05:03:41] [PASSED] None
[05:03:41] [PASSED] PAL
[05:03:41] [PASSED] NTSC
[05:03:41] [PASSED] Both, NTSC Default
[05:03:41] [PASSED] Both, PAL Default
[05:03:41] [PASSED] Both, NTSC Default, with PAL on command-line
[05:03:41] [PASSED] Both, PAL Default, with NTSC on command-line
[05:03:41] == [PASSED] drm_test_connector_helper_tv_get_modes_check ===
[05:03:41] ======== [PASSED] drm_connector_helper_tv_get_modes ========
[05:03:41] ================== drm_rect (9 subtests) ===================
[05:03:41] [PASSED] drm_test_rect_clip_scaled_div_by_zero
[05:03:41] [PASSED] drm_test_rect_clip_scaled_not_clipped
[05:03:41] [PASSED] drm_test_rect_clip_scaled_clipped
[05:03:41] [PASSED] drm_test_rect_clip_scaled_signed_vs_unsigned
[05:03:41] ================= drm_test_rect_intersect =================
[05:03:41] [PASSED] top-left x bottom-right: 2x2+1+1 x 2x2+0+0
[05:03:41] [PASSED] top-right x bottom-left: 2x2+0+0 x 2x2+1-1
[05:03:41] [PASSED] bottom-left x top-right: 2x2+1-1 x 2x2+0+0
[05:03:41] [PASSED] bottom-right x top-left: 2x2+0+0 x 2x2+1+1
[05:03:41] [PASSED] right x left: 2x1+0+0 x 3x1+1+0
[05:03:41] [PASSED] left x right: 3x1+1+0 x 2x1+0+0
[05:03:41] [PASSED] up x bottom: 1x2+0+0 x 1x3+0-1
[05:03:41] [PASSED] bottom x up: 1x3+0-1 x 1x2+0+0
[05:03:41] [PASSED] touching corner: 1x1+0+0 x 2x2+1+1
[05:03:41] [PASSED] touching side: 1x1+0+0 x 1x1+1+0
[05:03:41] [PASSED] equal rects: 2x2+0+0 x 2x2+0+0
[05:03:41] [PASSED] inside another: 2x2+0+0 x 1x1+1+1
[05:03:41] [PASSED] far away: 1x1+0+0 x 1x1+3+6
[05:03:41] [PASSED] points intersecting: 0x0+5+10 x 0x0+5+10
[05:03:41] [PASSED] points not intersecting: 0x0+0+0 x 0x0+5+10
[05:03:41] ============= [PASSED] drm_test_rect_intersect =============
[05:03:41] ================ drm_test_rect_calc_hscale ================
[05:03:41] [PASSED] normal use
[05:03:41] [PASSED] out of max range
[05:03:41] [PASSED] out of min range
[05:03:41] [PASSED] zero dst
[05:03:41] [PASSED] negative src
[05:03:41] [PASSED] negative dst
[05:03:41] ============ [PASSED] drm_test_rect_calc_hscale ============
[05:03:41] ================ drm_test_rect_calc_vscale ================
[05:03:41] [PASSED] normal use
[05:03:41] [PASSED] out of max range
[05:03:41] [PASSED] out of min range
[05:03:41] [PASSED] zero dst
[05:03:41] [PASSED] negative src
[05:03:41] [PASSED] negative dst
[05:03:41] ============ [PASSED] drm_test_rect_calc_vscale ============
[05:03:41] ================== drm_test_rect_rotate ===================
[05:03:41] [PASSED] reflect-x
[05:03:41] [PASSED] reflect-y
[05:03:41] [PASSED] rotate-0
[05:03:41] [PASSED] rotate-90
[05:03:41] [PASSED] rotate-180
[05:03:41] [PASSED] rotate-270
[05:03:41] ============== [PASSED] drm_test_rect_rotate ===============
[05:03:41] ================ drm_test_rect_rotate_inv =================
[05:03:41] [PASSED] reflect-x
[05:03:41] [PASSED] reflect-y
[05:03:41] [PASSED] rotate-0
[05:03:41] [PASSED] rotate-90
[05:03:41] [PASSED] rotate-180
[05:03:41] [PASSED] rotate-270
[05:03:41] ============ [PASSED] drm_test_rect_rotate_inv =============
[05:03:41] ==================== [PASSED] drm_rect =====================
[05:03:41] ============ drm_sysfb_modeset_test (1 subtest) ============
[05:03:41] ============ drm_test_sysfb_build_fourcc_list =============
[05:03:41] [PASSED] no native formats
[05:03:41] [PASSED] XRGB8888 as native format
[05:03:41] [PASSED] remove duplicates
[05:03:41] [PASSED] convert alpha formats
[05:03:41] [PASSED] random formats
[05:03:41] ======== [PASSED] drm_test_sysfb_build_fourcc_list =========
[05:03:41] ============= [PASSED] drm_sysfb_modeset_test ==============
[05:03:41] ================== drm_fixp (2 subtests) ===================
[05:03:41] [PASSED] drm_test_int2fixp
[05:03:41] [PASSED] drm_test_sm2fixp
[05:03:41] ==================== [PASSED] drm_fixp =====================
[05:03:41] ============================================================
[05:03:41] Testing complete. Ran 671 tests: passed: 671
[05:03:41] Elapsed time: 26.943s total, 1.794s configuring, 24.732s building, 0.412s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/ttm/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/ttm/tests/.kunitconfig
[05:03:41] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[05:03:43] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[05:03:53] Starting KUnit Kernel (1/1)...
[05:03:53] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[05:03:53] ============= refcount_interrupt (4 subtests) ==============
[05:03:53] [PASSED] test_single_irq_change
[05:03:53] [PASSED] test_nested_irq_change
[05:03:53] [PASSED] test_multiple_irq_change
[05:03:53] [PASSED] test_irq_save
[05:03:53] =============== [PASSED] refcount_interrupt ================
[05:03:53] ================= ttm_device (5 subtests) ==================
[05:03:53] [PASSED] ttm_device_init_basic
[05:03:53] [PASSED] ttm_device_init_multiple
[05:03:53] [PASSED] ttm_device_fini_basic
[05:03:53] [PASSED] ttm_device_init_no_vma_man
[05:03:53] ================== ttm_device_init_pools ==================
[05:03:53] [PASSED] No DMA allocations, no DMA32 required
[05:03:53] [PASSED] DMA allocations, DMA32 required
[05:03:53] [PASSED] No DMA allocations, DMA32 required
[05:03:53] [PASSED] DMA allocations, no DMA32 required
[05:03:53] ============== [PASSED] ttm_device_init_pools ==============
[05:03:53] =================== [PASSED] ttm_device ====================
[05:03:53] ================== ttm_pool (8 subtests) ===================
[05:03:53] ================== ttm_pool_alloc_basic ===================
[05:03:53] [PASSED] One page
[05:03:53] [PASSED] More than one page
[05:03:53] [PASSED] Above the allocation limit
[05:03:53] [PASSED] One page, with coherent DMA mappings enabled
[05:03:53] [PASSED] Above the allocation limit, with coherent DMA mappings enabled
[05:03:53] ============== [PASSED] ttm_pool_alloc_basic ===============
[05:03:53] ============== ttm_pool_alloc_basic_dma_addr ==============
[05:03:53] [PASSED] One page
[05:03:53] [PASSED] More than one page
[05:03:53] [PASSED] Above the allocation limit
[05:03:53] [PASSED] One page, with coherent DMA mappings enabled
[05:03:53] [PASSED] Above the allocation limit, with coherent DMA mappings enabled
[05:03:53] ========== [PASSED] ttm_pool_alloc_basic_dma_addr ==========
[05:03:53] [PASSED] ttm_pool_alloc_order_caching_match
[05:03:53] [PASSED] ttm_pool_alloc_caching_mismatch
[05:03:53] [PASSED] ttm_pool_alloc_order_mismatch
[05:03:53] [PASSED] ttm_pool_free_dma_alloc
[05:03:53] [PASSED] ttm_pool_free_no_dma_alloc
[05:03:53] [PASSED] ttm_pool_fini_basic
[05:03:53] ==================== [PASSED] ttm_pool =====================
[05:03:53] ================ ttm_resource (8 subtests) =================
[05:03:53] ================= ttm_resource_init_basic =================
[05:03:53] [PASSED] Init resource in TTM_PL_SYSTEM
[05:03:53] [PASSED] Init resource in TTM_PL_VRAM
[05:03:53] [PASSED] Init resource in a private placement
[05:03:53] [PASSED] Init resource in TTM_PL_SYSTEM, set placement flags
[05:03:53] ============= [PASSED] ttm_resource_init_basic =============
[05:03:53] [PASSED] ttm_resource_init_pinned
[05:03:53] [PASSED] ttm_resource_fini_basic
[05:03:53] [PASSED] ttm_resource_manager_init_basic
[05:03:53] [PASSED] ttm_resource_manager_usage_basic
[05:03:53] [PASSED] ttm_resource_manager_set_used_basic
[05:03:53] [PASSED] ttm_sys_man_alloc_basic
[05:03:53] [PASSED] ttm_sys_man_free_basic
[05:03:53] ================== [PASSED] ttm_resource ===================
[05:03:53] =================== ttm_tt (15 subtests) ===================
[05:03:53] ==================== ttm_tt_init_basic ====================
[05:03:53] [PASSED] Page-aligned size
[05:03:53] [PASSED] Extra pages requested
[05:03:53] ================ [PASSED] ttm_tt_init_basic ================
[05:03:53] [PASSED] ttm_tt_init_misaligned
[05:03:53] [PASSED] ttm_tt_fini_basic
[05:03:53] [PASSED] ttm_tt_fini_sg
[05:03:53] [PASSED] ttm_tt_fini_shmem
[05:03:53] [PASSED] ttm_tt_create_basic
[05:03:53] [PASSED] ttm_tt_create_invalid_bo_type
[05:03:53] [PASSED] ttm_tt_create_ttm_exists
[05:03:53] [PASSED] ttm_tt_create_failed
[05:03:53] [PASSED] ttm_tt_destroy_basic
[05:03:53] [PASSED] ttm_tt_populate_null_ttm
[05:03:53] [PASSED] ttm_tt_populate_populated_ttm
[05:03:53] [PASSED] ttm_tt_unpopulate_basic
[05:03:53] [PASSED] ttm_tt_unpopulate_empty_ttm
[05:03:53] [PASSED] ttm_tt_swapin_basic
[05:03:53] ===================== [PASSED] ttm_tt ======================
[05:03:53] =================== ttm_bo (14 subtests) ===================
[05:03:53] =========== ttm_bo_reserve_optimistic_no_ticket ===========
[05:03:53] [PASSED] Cannot be interrupted and sleeps
[05:03:53] [PASSED] Cannot be interrupted, locks straight away
[05:03:53] [PASSED] Can be interrupted, sleeps
[05:03:53] ======= [PASSED] ttm_bo_reserve_optimistic_no_ticket =======
[05:03:53] [PASSED] ttm_bo_reserve_locked_no_sleep
[05:03:53] [PASSED] ttm_bo_reserve_no_wait_ticket
[05:03:53] [PASSED] ttm_bo_reserve_double_resv
[05:03:53] [PASSED] ttm_bo_reserve_interrupted
[05:03:53] [PASSED] ttm_bo_reserve_deadlock
[05:03:53] [PASSED] ttm_bo_unreserve_basic
[05:03:53] [PASSED] ttm_bo_unreserve_pinned
[05:03:53] [PASSED] ttm_bo_unreserve_bulk
[05:03:53] [PASSED] ttm_bo_fini_basic
[05:03:53] [PASSED] ttm_bo_fini_shared_resv
[05:03:53] [PASSED] ttm_bo_pin_basic
[05:03:53] [PASSED] ttm_bo_pin_unpin_resource
[05:03:53] [PASSED] ttm_bo_multiple_pin_one_unpin
[05:03:53] ===================== [PASSED] ttm_bo ======================
[05:03:53] ============== ttm_bo_validate (22 subtests) ===============
[05:03:53] ============== ttm_bo_init_reserved_sys_man ===============
[05:03:53] [PASSED] Buffer object for userspace
[05:03:53] [PASSED] Kernel buffer object
[05:03:53] [PASSED] Shared buffer object
[05:03:53] ========== [PASSED] ttm_bo_init_reserved_sys_man ===========
[05:03:53] ============== ttm_bo_init_reserved_mock_man ==============
[05:03:53] [PASSED] Buffer object for userspace
[05:03:53] [PASSED] Kernel buffer object
[05:03:53] [PASSED] Shared buffer object
[05:03:53] ========== [PASSED] ttm_bo_init_reserved_mock_man ==========
[05:03:53] [PASSED] ttm_bo_init_reserved_resv
[05:03:53] ================== ttm_bo_validate_basic ==================
[05:03:53] [PASSED] Buffer object for userspace
[05:03:53] [PASSED] Kernel buffer object
[05:03:53] [PASSED] Shared buffer object
[05:03:53] ============== [PASSED] ttm_bo_validate_basic ==============
[05:03:53] [PASSED] ttm_bo_validate_invalid_placement
[05:03:53] ============= ttm_bo_validate_same_placement ==============
[05:03:53] [PASSED] System manager
[05:03:53] [PASSED] VRAM manager
[05:03:53] ========= [PASSED] ttm_bo_validate_same_placement ==========
[05:03:53] [PASSED] ttm_bo_validate_failed_alloc
[05:03:53] [PASSED] ttm_bo_validate_pinned
[05:03:53] [PASSED] ttm_bo_validate_busy_placement
[05:03:53] ================ ttm_bo_validate_multihop =================
[05:03:53] [PASSED] Buffer object for userspace
[05:03:53] [PASSED] Kernel buffer object
[05:03:53] [PASSED] Shared buffer object
[05:03:53] ============ [PASSED] ttm_bo_validate_multihop =============
[05:03:53] ========== ttm_bo_validate_no_placement_signaled ==========
[05:03:53] [PASSED] Buffer object in system domain, no page vector
[05:03:53] [PASSED] Buffer object in system domain with an existing page vector
[05:03:53] ====== [PASSED] ttm_bo_validate_no_placement_signaled ======
[05:03:53] ======== ttm_bo_validate_no_placement_not_signaled ========
[05:03:53] [PASSED] Buffer object for userspace
[05:03:53] [PASSED] Kernel buffer object
[05:03:53] [PASSED] Shared buffer object
[05:03:53] ==== [PASSED] ttm_bo_validate_no_placement_not_signaled ====
[05:03:53] [PASSED] ttm_bo_validate_move_fence_signaled
[05:03:53] ========= ttm_bo_validate_move_fence_not_signaled =========
[05:03:53] [PASSED] Waits for GPU
[05:03:53] [PASSED] Tries to lock straight away
[05:03:53] ===== [PASSED] ttm_bo_validate_move_fence_not_signaled =====
[05:03:53] [PASSED] ttm_bo_validate_swapout
[05:03:53] [PASSED] ttm_bo_validate_happy_evict
[05:03:53] [PASSED] ttm_bo_validate_all_pinned_evict
[05:03:53] [PASSED] ttm_bo_validate_allowed_only_evict
[05:03:53] [PASSED] ttm_bo_validate_deleted_evict
[05:03:53] [PASSED] ttm_bo_validate_busy_domain_evict
[05:03:53] [PASSED] ttm_bo_validate_evict_gutting
[05:03:53] [PASSED] ttm_bo_validate_recrusive_evict
[05:03:53] ================= [PASSED] ttm_bo_validate =================
[05:03:53] ============================================================
[05:03:53] Testing complete. Ran 106 tests: passed: 106
[05:03:53] Elapsed time: 12.099s total, 1.826s configuring, 10.058s building, 0.179s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/dma-buf/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/dma-buf/.kunitconfig
[05:03:54] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[05:03:55] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[05:04:04] Starting KUnit Kernel (1/1)...
[05:04:04] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[05:04:04] ============= refcount_interrupt (4 subtests) ==============
[05:04:04] [PASSED] test_single_irq_change
[05:04:04] [PASSED] test_nested_irq_change
[05:04:04] [PASSED] test_multiple_irq_change
[05:04:04] [PASSED] test_irq_save
[05:04:04] =============== [PASSED] refcount_interrupt ================
[05:04:04] =============== dma-buf-fence (12 subtests) ================
[05:04:04] [PASSED] test_sanitycheck
[05:04:04] [PASSED] test_signaling
[05:04:04] [PASSED] test_add_callback
[05:04:04] [PASSED] test_late_add_callback
[05:04:04] [PASSED] test_rm_callback
[05:04:04] [PASSED] test_late_rm_callback
[05:04:04] [PASSED] test_status
[05:04:04] [PASSED] test_error
[05:04:04] [PASSED] test_wait
[05:04:04] [PASSED] test_wait_timeout
[05:04:04] [PASSED] test_stub
[05:04:04] [SKIPPED] test_race_signal_callback (requires at least 2 CPUs)
[05:04:04] ================== [PASSED] dma-buf-fence ==================
[05:04:04] ============ dma-buf-fence-chain (11 subtests) =============
[05:04:04] [PASSED] test_sanitycheck
[05:04:04] [PASSED] test_find_seqno
[05:04:04] [PASSED] test_find_signaled
[05:04:04] [PASSED] test_find_out_of_order
[05:04:09] [PASSED] test_find_gap
[05:04:09] [PASSED] test_find_race
[05:04:09] [PASSED] test_signal_forward
[05:04:09] [PASSED] test_signal_backward
[05:04:09] [PASSED] test_wait_forward
[05:04:09] [PASSED] test_wait_backward
[05:04:09] [PASSED] test_wait_random
[05:04:09] =============== [PASSED] dma-buf-fence-chain ===============
[05:04:09] ============ dma-buf-fence-unwrap (10 subtests) ============
[05:04:09] [PASSED] test_sanitycheck
[05:04:09] [PASSED] test_unwrap_array
[05:04:09] [PASSED] test_unwrap_chain
[05:04:09] [PASSED] test_unwrap_chain_array
[05:04:09] [PASSED] test_unwrap_merge
[05:04:09] [PASSED] test_unwrap_merge_duplicate
[05:04:09] [PASSED] test_unwrap_merge_seqno
[05:04:09] [PASSED] test_unwrap_merge_order
[05:04:09] [PASSED] test_unwrap_merge_complex
[05:04:09] [PASSED] test_unwrap_merge_complex_seqno
[05:04:09] ============== [PASSED] dma-buf-fence-unwrap ===============
[05:04:09] ================ dma-buf-resv (5 subtests) =================
[05:04:09] [PASSED] test_sanitycheck
[05:04:09] ===================== test_signaling ======================
[05:04:09] [PASSED] kernel
[05:04:09] [PASSED] write
[05:04:09] [PASSED] read
[05:04:09] [PASSED] bookkeep
[05:04:09] ================= [PASSED] test_signaling ==================
[05:04:09] ====================== test_for_each ======================
[05:04:09] [PASSED] kernel
[05:04:09] [PASSED] write
[05:04:09] [PASSED] read
[05:04:09] [PASSED] bookkeep
[05:04:09] ================== [PASSED] test_for_each ==================
[05:04:09] ================= test_for_each_unlocked ==================
[05:04:09] [PASSED] kernel
[05:04:09] [PASSED] write
[05:04:09] [PASSED] read
[05:04:09] [PASSED] bookkeep
[05:04:09] ============= [PASSED] test_for_each_unlocked ==============
[05:04:09] ===================== test_get_fences =====================
[05:04:09] [PASSED] kernel
[05:04:09] [PASSED] write
[05:04:09] [PASSED] read
[05:04:09] [PASSED] bookkeep
[05:04:09] ================= [PASSED] test_get_fences =================
[05:04:09] ================== [PASSED] dma-buf-resv ===================
[05:04:09] ============================================================
[05:04:09] Testing complete. Ran 54 tests: passed: 53, skipped: 1
[05:04:09] Elapsed time: 15.866s total, 1.806s configuring, 8.689s building, 5.354s running
+ cleanup
++ stat -c %u:%g /kernel
+ chown -R 1003:1003 /kernel
^ permalink raw reply [flat|nested] 54+ messages in thread* ✓ Xe.CI.BAT: success for CPU binds and ULLS on migration queue (rev9)
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (25 preceding siblings ...)
2026-09-25 5:04 ` ✓ CI.KUnit: success " Patchwork
@ 2026-09-25 5:47 ` Patchwork
2026-09-25 15:05 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-25 18:08 ` [PATCH v7 00/24] CPU binds and ULLS on migration queue Maarten Lankhorst
28 siblings, 0 replies; 54+ messages in thread
From: Patchwork @ 2026-09-25 5:47 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
[-- Attachment #1: Type: text/plain, Size: 5953 bytes --]
== Series Details ==
Series: CPU binds and ULLS on migration queue (rev9)
URL : https://patchwork.freedesktop.org/series/149888/
State : success
== Summary ==
CI Bug Log - changes from xe-5824-999838292166407bbe911a41dee23112c49e9d44_BAT -> xe-pw-149888v9_BAT
====================================================
Summary
-------
**SUCCESS**
No regressions found.
Participating hosts (14 -> 15)
------------------------------
Additional (1): bat-wcl-2
Known issues
------------
Here are the changes found in xe-pw-149888v9_BAT that come from known issues:
### IGT changes ###
#### Issues hit ####
* igt@core_hotunplug@unbind-rebind:
- bat-bmg-2: [PASS][1] -> [ABORT][2] ([Intel XE#8007])
[1]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/bat-bmg-2/igt@core_hotunplug@unbind-rebind.html
[2]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-bmg-2/igt@core_hotunplug@unbind-rebind.html
* igt@fbdev@eof:
- bat-wcl-2: NOTRUN -> [SKIP][3] ([Intel XE#7241]) +4 other tests skip
[3]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@fbdev@eof.html
* igt@kms_addfb_basic@addfb25-y-tiled-small-legacy:
- bat-wcl-2: NOTRUN -> [SKIP][4] ([Intel XE#7245])
[4]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@kms_addfb_basic@addfb25-y-tiled-small-legacy.html
* igt@kms_flip@basic-flip-vs-wf_vblank:
- bat-wcl-2: NOTRUN -> [SKIP][5] ([Intel XE#7240]) +3 other tests skip
[5]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@kms_flip@basic-flip-vs-wf_vblank.html
* igt@kms_frontbuffer_tracking@basic:
- bat-wcl-2: NOTRUN -> [SKIP][6] ([Intel XE#7246])
[6]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@kms_frontbuffer_tracking@basic.html
* igt@kms_pipe_crc_basic@hang-read-crc:
- bat-wcl-2: NOTRUN -> [SKIP][7] ([Intel XE#7237]) +13 other tests skip
[7]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@kms_pipe_crc_basic@hang-read-crc.html
* igt@kms_psr@psr-sprite-plane-onoff:
- bat-wcl-2: NOTRUN -> [SKIP][8] ([Intel XE#2850]) +2 other tests skip
[8]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@kms_psr@psr-sprite-plane-onoff.html
* igt@xe_evict@evict-small-multi-vm:
- bat-wcl-2: NOTRUN -> [SKIP][9] ([Intel XE#7238]) +11 other tests skip
[9]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_evict@evict-small-multi-vm.html
* igt@xe_exec_balancer@twice-virtual-userptr-rebind:
- bat-wcl-2: NOTRUN -> [SKIP][10] ([Intel XE#7482]) +17 other tests skip
[10]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_exec_balancer@twice-virtual-userptr-rebind.html
* igt@xe_exec_multi_queue@many-queues-userptr-invalidate:
- bat-wcl-2: NOTRUN -> [SKIP][11] ([Intel XE#8364]) +13 other tests skip
[11]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_exec_multi_queue@many-queues-userptr-invalidate.html
* igt@xe_live_ktest@xe_bo@xe_bo_evict_kunit:
- bat-wcl-2: NOTRUN -> [SKIP][12] ([Intel XE#7239]) +2 other tests skip
[12]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_live_ktest@xe_bo@xe_bo_evict_kunit.html
* igt@xe_mmap@vram:
- bat-wcl-2: NOTRUN -> [SKIP][13] ([Intel XE#7243])
[13]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_mmap@vram.html
* igt@xe_pat@pat-index-xehpc:
- bat-wcl-2: NOTRUN -> [SKIP][14] ([Intel XE#7247] / [Intel XE#7590])
[14]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_pat@pat-index-xehpc.html
* igt@xe_pat@pat-index-xelp:
- bat-wcl-2: NOTRUN -> [SKIP][15] ([Intel XE#7242] / [Intel XE#7590])
[15]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_pat@pat-index-xelp.html
* igt@xe_pat@pat-index-xelpg:
- bat-wcl-2: NOTRUN -> [SKIP][16] ([Intel XE#7248] / [Intel XE#7590])
[16]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/bat-wcl-2/igt@xe_pat@pat-index-xelpg.html
[Intel XE#2850]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2850
[Intel XE#7237]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7237
[Intel XE#7238]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7238
[Intel XE#7239]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7239
[Intel XE#7240]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7240
[Intel XE#7241]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7241
[Intel XE#7242]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7242
[Intel XE#7243]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7243
[Intel XE#7245]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7245
[Intel XE#7246]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7246
[Intel XE#7247]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7247
[Intel XE#7248]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7248
[Intel XE#7482]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7482
[Intel XE#7590]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7590
[Intel XE#8007]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8007
[Intel XE#8364]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8364
Build changes
-------------
* Linux: xe-5824-999838292166407bbe911a41dee23112c49e9d44 -> xe-pw-149888v9
IGT_9112: 8d3284b5ae16ca1aef353dfaffb6c0896ceacc19 @ https://gitlab.freedesktop.org/drm/igt-gpu-tools.git
xe-5824-999838292166407bbe911a41dee23112c49e9d44: 999838292166407bbe911a41dee23112c49e9d44
xe-pw-149888v9: 149888v9
== Logs ==
For more details see: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/index.html
[-- Attachment #2: Type: text/html, Size: 6980 bytes --]
^ permalink raw reply [flat|nested] 54+ messages in thread* ✗ Xe.CI.FULL: failure for CPU binds and ULLS on migration queue (rev9)
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (26 preceding siblings ...)
2026-09-25 5:47 ` ✓ Xe.CI.BAT: " Patchwork
@ 2026-09-25 15:05 ` Patchwork
2026-09-25 16:01 ` Matthew Brost
2026-09-25 18:08 ` [PATCH v7 00/24] CPU binds and ULLS on migration queue Maarten Lankhorst
28 siblings, 1 reply; 54+ messages in thread
From: Patchwork @ 2026-09-25 15:05 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe
[-- Attachment #1: Type: text/plain, Size: 54862 bytes --]
== Series Details ==
Series: CPU binds and ULLS on migration queue (rev9)
URL : https://patchwork.freedesktop.org/series/149888/
State : failure
== Summary ==
CI Bug Log - changes from xe-5824-999838292166407bbe911a41dee23112c49e9d44_FULL -> xe-pw-149888v9_FULL
====================================================
Summary
-------
**FAILURE**
Serious unknown changes coming with xe-pw-149888v9_FULL absolutely need to be
verified manually.
If you think the reported changes have nothing to do with the changes
introduced in xe-pw-149888v9_FULL, please notify your bug team (I915-ci-infra@lists.freedesktop.org) to allow them
to document this new failure mode, which will reduce false positives in CI.
Participating hosts (3 -> 3)
------------------------------
No changes in participating hosts
Possible new issues
-------------------
Here are the unknown changes that may have been introduced in xe-pw-149888v9_FULL:
### IGT changes ###
#### Possible regressions ####
* igt@xe_vm@bind-array-enobufs:
- shard-ptl: [PASS][1] -> [FAIL][2]
[1]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-5/igt@xe_vm@bind-array-enobufs.html
[2]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-1/igt@xe_vm@bind-array-enobufs.html
- shard-lnl: [PASS][3] -> [FAIL][4]
[3]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-lnl-1/igt@xe_vm@bind-array-enobufs.html
[4]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-4/igt@xe_vm@bind-array-enobufs.html
- shard-bmg: [PASS][5] -> [FAIL][6]
[5]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-2/igt@xe_vm@bind-array-enobufs.html
[6]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-5/igt@xe_vm@bind-array-enobufs.html
Known issues
------------
Here are the changes found in xe-pw-149888v9_FULL that come from known issues:
### IGT changes ###
#### Issues hit ####
* igt@kms_async_flips@alternate-sync-async-flip-atomic:
- shard-bmg: [PASS][7] -> [FAIL][8] ([Intel XE#3718] / [Intel XE#6078])
[7]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-8/igt@kms_async_flips@alternate-sync-async-flip-atomic.html
[8]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_async_flips@alternate-sync-async-flip-atomic.html
* igt@kms_async_flips@alternate-sync-async-flip-atomic@pipe-b-dp-2:
- shard-bmg: [PASS][9] -> [FAIL][10] ([Intel XE#6078])
[9]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-8/igt@kms_async_flips@alternate-sync-async-flip-atomic@pipe-b-dp-2.html
[10]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_async_flips@alternate-sync-async-flip-atomic@pipe-b-dp-2.html
* igt@kms_async_flips@async-flip-suspend-resume@pipe-c-edp-1:
- shard-ptl: [PASS][11] -> [ABORT][12] ([Intel XE#9387])
[11]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-6/igt@kms_async_flips@async-flip-suspend-resume@pipe-c-edp-1.html
[12]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-1/igt@kms_async_flips@async-flip-suspend-resume@pipe-c-edp-1.html
* igt@kms_big_fb@4-tiled-max-hw-stride-64bpp-rotate-0-hflip:
- shard-lnl: NOTRUN -> [SKIP][13] ([Intel XE#1407])
[13]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_big_fb@4-tiled-max-hw-stride-64bpp-rotate-0-hflip.html
* igt@kms_big_fb@linear-64bpp-rotate-270:
- shard-bmg: NOTRUN -> [SKIP][14] ([Intel XE#2327]) +1 other test skip
[14]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_big_fb@linear-64bpp-rotate-270.html
* igt@kms_big_fb@y-tiled-addfb:
- shard-bmg: NOTRUN -> [SKIP][15] ([Intel XE#2328] / [Intel XE#7367])
[15]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_big_fb@y-tiled-addfb.html
* igt@kms_big_fb@yf-tiled-16bpp-rotate-270:
- shard-lnl: NOTRUN -> [SKIP][16] ([Intel XE#1124]) +1 other test skip
[16]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_big_fb@yf-tiled-16bpp-rotate-270.html
* igt@kms_big_fb@yf-tiled-max-hw-stride-32bpp-rotate-180-hflip-async-flip:
- shard-ptl: NOTRUN -> [SKIP][17] ([Intel XE#5808]) +1 other test skip
[17]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@kms_big_fb@yf-tiled-max-hw-stride-32bpp-rotate-180-hflip-async-flip.html
* igt@kms_big_fb@yf-tiled-max-hw-stride-64bpp-rotate-0-hflip:
- shard-bmg: NOTRUN -> [SKIP][18] ([Intel XE#1124]) +4 other tests skip
[18]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_big_fb@yf-tiled-max-hw-stride-64bpp-rotate-0-hflip.html
* igt@kms_bw@connected-linear-tiling-2-displays-target-3840x2160p:
- shard-ptl: NOTRUN -> [SKIP][19] ([Intel XE#7679])
[19]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_bw@connected-linear-tiling-2-displays-target-3840x2160p.html
- shard-lnl: NOTRUN -> [SKIP][20] ([Intel XE#7679])
[20]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_bw@connected-linear-tiling-2-displays-target-3840x2160p.html
* igt@kms_bw@connected-linear-tiling-3-displays-target-1920x1080p:
- shard-bmg: NOTRUN -> [SKIP][21] ([Intel XE#7679])
[21]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_bw@connected-linear-tiling-3-displays-target-1920x1080p.html
* igt@kms_ccs@crc-primary-rotation-180-y-tiled-gen12-rc-ccs:
- shard-ptl: NOTRUN -> [SKIP][22] ([Intel XE#5822]) +4 other tests skip
[22]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_ccs@crc-primary-rotation-180-y-tiled-gen12-rc-ccs.html
- shard-lnl: NOTRUN -> [SKIP][23] ([Intel XE#2887]) +3 other tests skip
[23]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_ccs@crc-primary-rotation-180-y-tiled-gen12-rc-ccs.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-a-hdmi-a-3:
- shard-bmg: [PASS][24] -> [ABORT][25] ([Intel XE#9362])
[24]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-7/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-a-hdmi-a-3.html
[25]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-a-hdmi-a-3.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-d-hdmi-a-3:
- shard-bmg: NOTRUN -> [INCOMPLETE][26] ([Intel XE#7084] / [Intel XE#8150])
[26]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-d-hdmi-a-3.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-lnl-ccs@pipe-d-edp-1:
- shard-ptl: NOTRUN -> [ABORT][27] ([Intel XE#9367] / [Intel XE#9375]) +1 other test abort
[27]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_ccs@crc-primary-suspend-4-tiled-lnl-ccs@pipe-d-edp-1.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-mtl-mc-ccs:
- shard-bmg: NOTRUN -> [SKIP][28] ([Intel XE#3432])
[28]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_ccs@crc-primary-suspend-4-tiled-mtl-mc-ccs.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-mtl-rc-ccs:
- shard-lnl: NOTRUN -> [SKIP][29] ([Intel XE#3432])
[29]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_ccs@crc-primary-suspend-4-tiled-mtl-rc-ccs.html
* igt@kms_ccs@crc-sprite-planes-basic-4-tiled-dg2-rc-ccs-cc:
- shard-bmg: NOTRUN -> [SKIP][30] ([Intel XE#2887]) +7 other tests skip
[30]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_ccs@crc-sprite-planes-basic-4-tiled-dg2-rc-ccs-cc.html
* igt@kms_chamelium_color@ctm-max:
- shard-ptl: NOTRUN -> [SKIP][31] ([Intel XE#5824])
[31]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_chamelium_color@ctm-max.html
- shard-lnl: NOTRUN -> [SKIP][32] ([Intel XE#306] / [Intel XE#7358])
[32]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_chamelium_color@ctm-max.html
* igt@kms_chamelium_color_pipeline@plane-lut1d-pre-ctm3x4:
- shard-bmg: NOTRUN -> [SKIP][33] ([Intel XE#7358]) +1 other test skip
[33]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_chamelium_color_pipeline@plane-lut1d-pre-ctm3x4.html
* igt@kms_chamelium_hpd@dp-hpd-after-hibernate:
- shard-bmg: NOTRUN -> [SKIP][34] ([Intel XE#2252]) +2 other tests skip
[34]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_chamelium_hpd@dp-hpd-after-hibernate.html
* igt@kms_chamelium_hpd@hdmi-hpd-fast:
- shard-ptl: NOTRUN -> [SKIP][35] ([Intel XE#2252])
[35]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_chamelium_hpd@hdmi-hpd-fast.html
- shard-lnl: NOTRUN -> [SKIP][36] ([Intel XE#373]) +2 other tests skip
[36]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_chamelium_hpd@hdmi-hpd-fast.html
* igt@kms_color_pipeline@plane-fixed-matrix-yuv-rgb-bt601-lim-lut1d@pipe-b-plane-0:
- shard-bmg: NOTRUN -> [FAIL][37] ([Intel XE#9333]) +9 other tests fail
[37]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_color_pipeline@plane-fixed-matrix-yuv-rgb-bt601-lim-lut1d@pipe-b-plane-0.html
* igt@kms_color_pipeline@plane-lut3d-green-only:
- shard-lnl: NOTRUN -> [SKIP][38] ([Intel XE#6969] / [Intel XE#7006]) +1 other test skip
[38]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_color_pipeline@plane-lut3d-green-only.html
* igt@kms_color_pipeline@plane-lut3d-green-only@pipe-b-plane-4:
- shard-ptl: NOTRUN -> [SKIP][39] ([Intel XE#6969]) +20 other tests skip
[39]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_color_pipeline@plane-lut3d-green-only@pipe-b-plane-4.html
- shard-lnl: NOTRUN -> [SKIP][40] ([Intel XE#6969]) +13 other tests skip
[40]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_color_pipeline@plane-lut3d-green-only@pipe-b-plane-4.html
* igt@kms_content_protection@atomic:
- shard-bmg: NOTRUN -> [FAIL][41] ([Intel XE#1178] / [Intel XE#3304] / [Intel XE#7374]) +3 other tests fail
[41]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_content_protection@atomic.html
* igt@kms_content_protection@legacy:
- shard-ptl: NOTRUN -> [SKIP][42] ([Intel XE#7642])
[42]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_content_protection@legacy.html
- shard-lnl: NOTRUN -> [SKIP][43] ([Intel XE#7642])
[43]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_content_protection@legacy.html
* igt@kms_cursor_crc@cursor-offscreen-128x42:
- shard-lnl: NOTRUN -> [SKIP][44] ([Intel XE#1424])
[44]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_cursor_crc@cursor-offscreen-128x42.html
* igt@kms_cursor_crc@cursor-onscreen-32x10:
- shard-bmg: NOTRUN -> [SKIP][45] ([Intel XE#2320])
[45]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_cursor_crc@cursor-onscreen-32x10.html
* igt@kms_cursor_crc@cursor-random-512x512:
- shard-lnl: NOTRUN -> [SKIP][46] ([Intel XE#2321] / [Intel XE#7355])
[46]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_cursor_crc@cursor-random-512x512.html
* igt@kms_cursor_crc@cursor-rapid-movement-512x170:
- shard-bmg: NOTRUN -> [SKIP][47] ([Intel XE#2321] / [Intel XE#7355])
[47]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_cursor_crc@cursor-rapid-movement-512x170.html
* igt@kms_cursor_crc@cursor-sliding-max-size:
- shard-ptl: NOTRUN -> [SKIP][48] ([Intel XE#5900]) +1 other test skip
[48]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@kms_cursor_crc@cursor-sliding-max-size.html
* igt@kms_cursor_crc@cursor-suspend@pipe-d-edp-1:
- shard-ptl: NOTRUN -> [ABORT][49] ([Intel XE#9375])
[49]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-8/igt@kms_cursor_crc@cursor-suspend@pipe-d-edp-1.html
* igt@kms_cursor_legacy@short-busy-flip-before-cursor-toggle:
- shard-ptl: NOTRUN -> [SKIP][50] ([Intel XE#6035])
[50]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_cursor_legacy@short-busy-flip-before-cursor-toggle.html
- shard-lnl: NOTRUN -> [SKIP][51] ([Intel XE#323] / [Intel XE#6035])
[51]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_cursor_legacy@short-busy-flip-before-cursor-toggle.html
* igt@kms_dsc@dsc-with-bpc-ultrajoiner:
- shard-bmg: NOTRUN -> [SKIP][52] ([Intel XE#8265])
[52]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_dsc@dsc-with-bpc-ultrajoiner.html
* igt@kms_dsc@dsc-with-output-formats:
- shard-ptl: NOTRUN -> [SKIP][53] ([Intel XE#5907])
[53]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_dsc@dsc-with-output-formats.html
- shard-lnl: NOTRUN -> [SKIP][54] ([Intel XE#8265])
[54]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_dsc@dsc-with-output-formats.html
* igt@kms_fbcon_fbt@fbc:
- shard-bmg: NOTRUN -> [SKIP][55] ([Intel XE#4156] / [Intel XE#7425])
[55]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_fbcon_fbt@fbc.html
* igt@kms_feature_discovery@display-3x:
- shard-bmg: NOTRUN -> [SKIP][56] ([Intel XE#2373] / [Intel XE#7448])
[56]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_feature_discovery@display-3x.html
* igt@kms_flip@2x-dpms-vs-vblank-race:
- shard-lnl: NOTRUN -> [SKIP][57] ([Intel XE#1421]) +2 other tests skip
[57]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_flip@2x-dpms-vs-vblank-race.html
* igt@kms_flip@2x-flip-vs-blocking-wf-vblank:
- shard-ptl: NOTRUN -> [SKIP][58] ([Intel XE#2316]) +2 other tests skip
[58]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_flip@2x-flip-vs-blocking-wf-vblank.html
* igt@kms_flip@basic-plain-flip@d-hdmi-a3:
- shard-bmg: [PASS][59] -> [DMESG-WARN][60] ([Intel XE#1727] / [Intel XE#6819]) +1 other test dmesg-warn
[59]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-6/igt@kms_flip@basic-plain-flip@d-hdmi-a3.html
[60]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-3/igt@kms_flip@basic-plain-flip@d-hdmi-a3.html
* igt@kms_flip@flip-vs-expired-vblank-interruptible@c-edp1:
- shard-lnl: [PASS][61] -> [FAIL][62] ([Intel XE#301] / [Intel XE#3149]) +1 other test fail
[61]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-lnl-8/igt@kms_flip@flip-vs-expired-vblank-interruptible@c-edp1.html
[62]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-2/igt@kms_flip@flip-vs-expired-vblank-interruptible@c-edp1.html
* igt@kms_flip_scaled_crc@flip-32bpp-yftile-to-32bpp-yftileccs-upscaling:
- shard-ptl: NOTRUN -> [SKIP][63] ([Intel XE#7178]) +1 other test skip
[63]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_flip_scaled_crc@flip-32bpp-yftile-to-32bpp-yftileccs-upscaling.html
- shard-lnl: NOTRUN -> [SKIP][64] ([Intel XE#7178] / [Intel XE#7351])
[64]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_flip_scaled_crc@flip-32bpp-yftile-to-32bpp-yftileccs-upscaling.html
* igt@kms_flip_scaled_crc@flip-32bpp-yftileccs-to-64bpp-yftile-upscaling:
- shard-bmg: NOTRUN -> [SKIP][65] ([Intel XE#7178] / [Intel XE#7351]) +1 other test skip
[65]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_flip_scaled_crc@flip-32bpp-yftileccs-to-64bpp-yftile-upscaling.html
* igt@kms_flip_scaled_crc@flip-64bpp-linear-to-32bpp-linear-downscaling:
- shard-lnl: NOTRUN -> [SKIP][66] ([Intel XE#1397] / [Intel XE#1745] / [Intel XE#7385] / [Intel XE#9144])
[66]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_flip_scaled_crc@flip-64bpp-linear-to-32bpp-linear-downscaling.html
* igt@kms_flip_scaled_crc@flip-64bpp-linear-to-32bpp-linear-downscaling@pipe-a-default-mode:
- shard-lnl: NOTRUN -> [SKIP][67] ([Intel XE#1397] / [Intel XE#7385] / [Intel XE#9144])
[67]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_flip_scaled_crc@flip-64bpp-linear-to-32bpp-linear-downscaling@pipe-a-default-mode.html
* igt@kms_frontbuffer_tracking@drrs-1p-pri-indfb-multidraw:
- shard-lnl: NOTRUN -> [SKIP][68] ([Intel XE#6312] / [Intel XE#651]) +6 other tests skip
[68]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_frontbuffer_tracking@drrs-1p-pri-indfb-multidraw.html
* igt@kms_frontbuffer_tracking@drrs-1p-primscrn-indfb-plflip-blt:
- shard-ptl: NOTRUN -> [SKIP][69] ([Intel XE#5800] / [Intel XE#6312]) +4 other tests skip
[69]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_frontbuffer_tracking@drrs-1p-primscrn-indfb-plflip-blt.html
* igt@kms_frontbuffer_tracking@drrshdr-argb161616f-draw-mmap-wc:
- shard-bmg: NOTRUN -> [SKIP][70] ([Intel XE#7061]) +3 other tests skip
[70]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_frontbuffer_tracking@drrshdr-argb161616f-draw-mmap-wc.html
* igt@kms_frontbuffer_tracking@fbc-1p-pri-indfb-multidraw:
- shard-bmg: NOTRUN -> [SKIP][71] ([Intel XE#4141]) +6 other tests skip
[71]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_frontbuffer_tracking@fbc-1p-pri-indfb-multidraw.html
* igt@kms_frontbuffer_tracking@fbc-abgr161616f-draw-render:
- shard-ptl: NOTRUN -> [SKIP][72] ([Intel XE#7061] / [Intel XE#7356])
[72]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@kms_frontbuffer_tracking@fbc-abgr161616f-draw-render.html
* igt@kms_frontbuffer_tracking@fbcdrrs-2p-scndscrn-pri-indfb-draw-render:
- shard-lnl: NOTRUN -> [SKIP][73] ([Intel XE#656] / [Intel XE#7905]) +7 other tests skip
[73]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_frontbuffer_tracking@fbcdrrs-2p-scndscrn-pri-indfb-draw-render.html
* igt@kms_frontbuffer_tracking@fbcdrrs-abgr161616f-draw-mmap-wc:
- shard-bmg: NOTRUN -> [SKIP][74] ([Intel XE#7061] / [Intel XE#7356]) +1 other test skip
[74]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_frontbuffer_tracking@fbcdrrs-abgr161616f-draw-mmap-wc.html
* igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-offscreen-pri-shrfb-draw-blt:
- shard-lnl: NOTRUN -> [SKIP][75] ([Intel XE#6312]) +1 other test skip
[75]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-offscreen-pri-shrfb-draw-blt.html
* igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-primscrn-pri-indfb-draw-blt:
- shard-ptl: NOTRUN -> [SKIP][76] ([Intel XE#6312]) +1 other test skip
[76]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-primscrn-pri-indfb-draw-blt.html
* igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-primscrn-pri-shrfb-draw-render:
- shard-bmg: NOTRUN -> [SKIP][77] ([Intel XE#2311]) +24 other tests skip
[77]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_frontbuffer_tracking@fbcdrrshdr-1p-primscrn-pri-shrfb-draw-render.html
* igt@kms_frontbuffer_tracking@fbcpsrhdr-1p-primscrn-spr-indfb-draw-mmap-wc:
- shard-bmg: NOTRUN -> [SKIP][78] ([Intel XE#2313]) +24 other tests skip
[78]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@kms_frontbuffer_tracking@fbcpsrhdr-1p-primscrn-spr-indfb-draw-mmap-wc.html
* igt@kms_frontbuffer_tracking@fbcpsrhdr-abgr161616f-draw-render:
- shard-ptl: NOTRUN -> [SKIP][79] ([Intel XE#7061]) +1 other test skip
[79]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_frontbuffer_tracking@fbcpsrhdr-abgr161616f-draw-render.html
- shard-lnl: NOTRUN -> [SKIP][80] ([Intel XE#7061]) +1 other test skip
[80]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_frontbuffer_tracking@fbcpsrhdr-abgr161616f-draw-render.html
* igt@kms_frontbuffer_tracking@hdr-2p-scndscrn-pri-shrfb-draw-mmap-wc:
- shard-ptl: NOTRUN -> [SKIP][81] ([Intel XE#5812]) +18 other tests skip
[81]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_frontbuffer_tracking@hdr-2p-scndscrn-pri-shrfb-draw-mmap-wc.html
- shard-lnl: NOTRUN -> [SKIP][82] ([Intel XE#7905]) +7 other tests skip
[82]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_frontbuffer_tracking@hdr-2p-scndscrn-pri-shrfb-draw-mmap-wc.html
* igt@kms_frontbuffer_tracking@psrhdr-1p-primscrn-indfb-plflip-blt:
- shard-lnl: NOTRUN -> [SKIP][83] ([Intel XE#7865]) +6 other tests skip
[83]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_frontbuffer_tracking@psrhdr-1p-primscrn-indfb-plflip-blt.html
* igt@kms_frontbuffer_tracking@psrhdr-1p-primscrn-spr-indfb-draw-blt:
- shard-ptl: NOTRUN -> [SKIP][84] ([Intel XE#7865]) +4 other tests skip
[84]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_frontbuffer_tracking@psrhdr-1p-primscrn-spr-indfb-draw-blt.html
* igt@kms_joiner@basic-max-non-joiner:
- shard-lnl: NOTRUN -> [SKIP][85] ([Intel XE#4298] / [Intel XE#5873])
[85]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_joiner@basic-max-non-joiner.html
* igt@kms_plane@pixel-format-4-tiled-mtl-rc-ccs-cc-modifier:
- shard-lnl: NOTRUN -> [SKIP][86] ([Intel XE#7283])
[86]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_plane@pixel-format-4-tiled-mtl-rc-ccs-cc-modifier.html
* igt@kms_plane@plane-panning-bottom-right-suspend:
- shard-ptl: [PASS][87] -> [ABORT][88] ([Intel XE#9362]) +6 other tests abort
[87]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-3/igt@kms_plane@plane-panning-bottom-right-suspend.html
[88]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-7/igt@kms_plane@plane-panning-bottom-right-suspend.html
* igt@kms_plane_scaling@planes-downscale-factor-0-5@pipe-c:
- shard-bmg: NOTRUN -> [SKIP][89] ([Intel XE#2763] / [Intel XE#6886]) +4 other tests skip
[89]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_plane_scaling@planes-downscale-factor-0-5@pipe-c.html
* igt@kms_pm_dc@dc5-pageflip-negative:
- shard-bmg: NOTRUN -> [SKIP][90] ([Intel XE#6927] / [Intel XE#8854])
[90]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_pm_dc@dc5-pageflip-negative.html
* igt@kms_pm_rpm@dpms-mode-unset-non-lpsp:
- shard-ptl: NOTRUN -> [SKIP][91] ([Intel XE#5810])
[91]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_pm_rpm@dpms-mode-unset-non-lpsp.html
- shard-lnl: NOTRUN -> [SKIP][92] ([Intel XE#1439] / [Intel XE#7402] / [Intel XE#836])
[92]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_pm_rpm@dpms-mode-unset-non-lpsp.html
* igt@kms_psr2_sf@pr-cursor-plane-move-continuous-exceed-sf:
- shard-bmg: NOTRUN -> [SKIP][93] ([Intel XE#1489]) +2 other tests skip
[93]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_psr2_sf@pr-cursor-plane-move-continuous-exceed-sf.html
* igt@kms_psr2_sf@pr-plane-move-sf-dmg-area:
- shard-ptl: NOTRUN -> [SKIP][94] ([Intel XE#7304])
[94]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_psr2_sf@pr-plane-move-sf-dmg-area.html
- shard-lnl: NOTRUN -> [SKIP][95] ([Intel XE#2893] / [Intel XE#7304])
[95]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@kms_psr2_sf@pr-plane-move-sf-dmg-area.html
* igt@kms_psr2_su@page_flip-p010:
- shard-bmg: NOTRUN -> [SKIP][96] ([Intel XE#2387] / [Intel XE#7429])
[96]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_psr2_su@page_flip-p010.html
* igt@kms_psr@pr-basic:
- shard-bmg: NOTRUN -> [SKIP][97] ([Intel XE#2234] / [Intel XE#2850])
[97]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@kms_psr@pr-basic.html
* igt@kms_psr@pr-cursor-plane-onoff:
- shard-lnl: NOTRUN -> [SKIP][98] ([Intel XE#1406])
[98]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@kms_psr@pr-cursor-plane-onoff.html
* igt@kms_psr@pr-sprite-render:
- shard-ptl: NOTRUN -> [SKIP][99] ([Intel XE#1406])
[99]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@kms_psr@pr-sprite-render.html
* igt@kms_vblank@ts-continuation-suspend@pipe-c-edp-1:
- shard-lnl: [PASS][100] -> [ABORT][101] ([Intel XE#9396]) +1 other test abort
[100]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-lnl-4/igt@kms_vblank@ts-continuation-suspend@pipe-c-edp-1.html
[101]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-2/igt@kms_vblank@ts-continuation-suspend@pipe-c-edp-1.html
* igt@kms_vrr@flip-suspend@pipe-a-edp-1:
- shard-ptl: [PASS][102] -> [FAIL][103] ([Intel XE#4227] / [Intel XE#7397]) +1 other test fail
[102]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-7/igt@kms_vrr@flip-suspend@pipe-a-edp-1.html
[103]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@kms_vrr@flip-suspend@pipe-a-edp-1.html
* igt@xe_evict@evict-mixed-many-threads-small:
- shard-lnl: NOTRUN -> [SKIP][104] ([Intel XE#6540] / [Intel XE#688]) +3 other tests skip
[104]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_evict@evict-mixed-many-threads-small.html
* igt@xe_evict@evict-mixed-threads-large-multi-vm:
- shard-ptl: NOTRUN -> [SKIP][105] ([Intel XE#5764]) +3 other tests skip
[105]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@xe_evict@evict-mixed-threads-large-multi-vm.html
* igt@xe_exec_balancer@many-virtual-userptr-invalidate-race:
- shard-lnl: NOTRUN -> [SKIP][106] ([Intel XE#7482]) +2 other tests skip
[106]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_exec_balancer@many-virtual-userptr-invalidate-race.html
* igt@xe_exec_balancer@twice-virtual-rebind:
- shard-ptl: NOTRUN -> [SKIP][107] ([Intel XE#7482]) +5 other tests skip
[107]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@xe_exec_balancer@twice-virtual-rebind.html
* igt@xe_exec_basic@multigpu-many-execqueues-many-vm-bindexecqueue-userptr-rebind:
- shard-ptl: NOTRUN -> [SKIP][108] ([Intel XE#5575] / [Intel XE#7484]) +1 other test skip
[108]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@xe_exec_basic@multigpu-many-execqueues-many-vm-bindexecqueue-userptr-rebind.html
- shard-lnl: NOTRUN -> [SKIP][109] ([Intel XE#1392]) +1 other test skip
[109]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_exec_basic@multigpu-many-execqueues-many-vm-bindexecqueue-userptr-rebind.html
* igt@xe_exec_basic@multigpu-no-exec-bindexecqueue-rebind:
- shard-bmg: NOTRUN -> [SKIP][110] ([Intel XE#2322] / [Intel XE#7372]) +4 other tests skip
[110]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@xe_exec_basic@multigpu-no-exec-bindexecqueue-rebind.html
* igt@xe_exec_fault_mode@many-execqueues-multi-queue-imm:
- shard-bmg: NOTRUN -> [SKIP][111] ([Intel XE#8374]) +6 other tests skip
[111]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@xe_exec_fault_mode@many-execqueues-multi-queue-imm.html
* igt@xe_exec_fault_mode@once-multi-queue-userptr:
- shard-lnl: NOTRUN -> [SKIP][112] ([Intel XE#8374])
[112]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_exec_fault_mode@once-multi-queue-userptr.html
* igt@xe_exec_fault_mode@twice-multi-queue:
- shard-ptl: NOTRUN -> [SKIP][113] ([Intel XE#8374]) +1 other test skip
[113]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@xe_exec_fault_mode@twice-multi-queue.html
* igt@xe_exec_multi_queue@many-execs-preempt-mode-close-fd:
- shard-bmg: NOTRUN -> [SKIP][114] ([Intel XE#8364]) +12 other tests skip
[114]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@xe_exec_multi_queue@many-execs-preempt-mode-close-fd.html
* igt@xe_exec_multi_queue@one-queue-close-fd-smem:
- shard-ptl: NOTRUN -> [SKIP][115] ([Intel XE#8364]) +7 other tests skip
[115]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@xe_exec_multi_queue@one-queue-close-fd-smem.html
* igt@xe_exec_multi_queue@two-queues-close-fd-smem:
- shard-lnl: NOTRUN -> [SKIP][116] ([Intel XE#8364]) +12 other tests skip
[116]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_exec_multi_queue@two-queues-close-fd-smem.html
* igt@xe_exec_reset@cm-multi-queue-close-execqueues:
- shard-ptl: NOTRUN -> [SKIP][117] ([Intel XE#8369])
[117]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@xe_exec_reset@cm-multi-queue-close-execqueues.html
* igt@xe_exec_threads@threads-hang-userptr-invalidate-race:
- shard-bmg: [PASS][118] -> [FAIL][119] ([Intel XE#9096])
[118]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-10/igt@xe_exec_threads@threads-hang-userptr-invalidate-race.html
[119]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-4/igt@xe_exec_threads@threads-hang-userptr-invalidate-race.html
* igt@xe_exec_threads@threads-multi-queue-mixed-fd-userptr-rebind:
- shard-bmg: NOTRUN -> [SKIP][120] ([Intel XE#8378]) +3 other tests skip
[120]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@xe_exec_threads@threads-multi-queue-mixed-fd-userptr-rebind.html
* igt@xe_exec_threads@threads-multi-queue-mixed-rebind:
- shard-lnl: NOTRUN -> [SKIP][121] ([Intel XE#8378])
[121]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@xe_exec_threads@threads-multi-queue-mixed-rebind.html
* igt@xe_fault_injection@exec-queue-create-fail-xe_vm_add_compute_exec_queue:
- shard-bmg: [PASS][122] -> [ABORT][123] ([Intel XE#8007])
[122]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-3/igt@xe_fault_injection@exec-queue-create-fail-xe_vm_add_compute_exec_queue.html
[123]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-4/igt@xe_fault_injection@exec-queue-create-fail-xe_vm_add_compute_exec_queue.html
* igt@xe_multigpu_svm@mgpu-atomic-op-basic:
- shard-bmg: NOTRUN -> [SKIP][124] ([Intel XE#6964]) +2 other tests skip
[124]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@xe_multigpu_svm@mgpu-atomic-op-basic.html
* igt@xe_multigpu_svm@mgpu-atomic-op-conflict:
- shard-ptl: NOTRUN -> [SKIP][125] ([Intel XE#6964])
[125]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@xe_multigpu_svm@mgpu-atomic-op-conflict.html
- shard-lnl: NOTRUN -> [SKIP][126] ([Intel XE#6964])
[126]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_multigpu_svm@mgpu-atomic-op-conflict.html
* igt@xe_page_reclaim@binds-large-split:
- shard-bmg: NOTRUN -> [SKIP][127] ([Intel XE#7793])
[127]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@xe_page_reclaim@binds-large-split.html
* igt@xe_pat@l2-flush-opt-svm-pat-restrict:
- shard-lnl: NOTRUN -> [SKIP][128] ([Intel XE#7590] / [Intel XE#7772])
[128]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@xe_pat@l2-flush-opt-svm-pat-restrict.html
* igt@xe_pm@d3cold-mmap-system:
- shard-bmg: NOTRUN -> [SKIP][129] ([Intel XE#2284] / [Intel XE#7370])
[129]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@xe_pm@d3cold-mmap-system.html
* igt@xe_pm@s3-basic-exec:
- shard-ptl: NOTRUN -> [SKIP][130] ([Intel XE#6867])
[130]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@xe_pm@s3-basic-exec.html
- shard-lnl: NOTRUN -> [SKIP][131] ([Intel XE#584] / [Intel XE#7369])
[131]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_pm@s3-basic-exec.html
* igt@xe_pm@s4-basic-exec:
- shard-bmg: [PASS][132] -> [ABORT][133] ([Intel XE#9399])
[132]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-7/igt@xe_pm@s4-basic-exec.html
[133]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@xe_pm@s4-basic-exec.html
* igt@xe_pmu@engine-activity-suspend:
- shard-ptl: NOTRUN -> [ABORT][134] ([Intel XE#9362])
[134]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@xe_pmu@engine-activity-suspend.html
* igt@xe_prefetch_fault@prefetch-fault-svm:
- shard-ptl: NOTRUN -> [SKIP][135] ([Intel XE#7599])
[135]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@xe_prefetch_fault@prefetch-fault-svm.html
- shard-lnl: NOTRUN -> [SKIP][136] ([Intel XE#7599])
[136]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_prefetch_fault@prefetch-fault-svm.html
* igt@xe_pxp@pxp-termination-key-update-post-rpm:
- shard-bmg: NOTRUN -> [SKIP][137] ([Intel XE#4733] / [Intel XE#7417])
[137]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@xe_pxp@pxp-termination-key-update-post-rpm.html
* igt@xe_query@multigpu-query-invalid-cs-cycles:
- shard-bmg: NOTRUN -> [SKIP][138] ([Intel XE#944])
[138]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@xe_query@multigpu-query-invalid-cs-cycles.html
* igt@xe_sriov_scheduling@equal-throughput-low-priority:
- shard-lnl: NOTRUN -> [SKIP][139] ([Intel XE#8339])
[139]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@xe_sriov_scheduling@equal-throughput-low-priority.html
* igt@xe_sriov_vfio@region-info:
- shard-lnl: NOTRUN -> [SKIP][140] ([Intel XE#7724])
[140]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_sriov_vfio@region-info.html
* igt@xe_wedged@basic-wedged:
- shard-ptl: NOTRUN -> [DMESG-WARN][141] ([Intel XE#8963])
[141]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-6/igt@xe_wedged@basic-wedged.html
- shard-lnl: NOTRUN -> [DMESG-WARN][142] ([Intel XE#8963])
[142]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-8/igt@xe_wedged@basic-wedged.html
#### Possible fixes ####
* igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-b-hdmi-a-3:
- shard-bmg: [INCOMPLETE][143] ([Intel XE#7084] / [Intel XE#8150]) -> [PASS][144]
[143]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-7/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-b-hdmi-a-3.html
[144]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs@pipe-b-hdmi-a-3.html
* igt@kms_cursor_crc@cursor-suspend@pipe-a-edp-1:
- shard-ptl: [ABORT][145] ([Intel XE#9362]) -> [PASS][146] +3 other tests pass
[145]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-6/igt@kms_cursor_crc@cursor-suspend@pipe-a-edp-1.html
[146]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-8/igt@kms_cursor_crc@cursor-suspend@pipe-a-edp-1.html
* igt@kms_hdr@static-toggle-suspend:
- shard-bmg: [ABORT][147] -> [PASS][148] +1 other test pass
[147]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-7/igt@kms_hdr@static-toggle-suspend.html
[148]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-4/igt@kms_hdr@static-toggle-suspend.html
* igt@kms_vblank@accuracy-idle:
- shard-ptl: [FAIL][149] -> [PASS][150] +1 other test pass
[149]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-6/igt@kms_vblank@accuracy-idle.html
[150]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-3/igt@kms_vblank@accuracy-idle.html
* igt@kms_vblank@ts-continuation-suspend:
- shard-ptl: [ABORT][151] ([Intel XE#9375]) -> [PASS][152] +1 other test pass
[151]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-8/igt@kms_vblank@ts-continuation-suspend.html
[152]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-5/igt@kms_vblank@ts-continuation-suspend.html
* igt@kms_vrr@flip-dpms@pipe-a-edp-1:
- shard-ptl: [FAIL][153] ([Intel XE#4227] / [Intel XE#7397]) -> [PASS][154] +3 other tests pass
[153]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-7/igt@kms_vrr@flip-dpms@pipe-a-edp-1.html
[154]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-2/igt@kms_vrr@flip-dpms@pipe-a-edp-1.html
* igt@xe_evict@evict-mixed-many-threads-small:
- shard-bmg: [INCOMPLETE][155] ([Intel XE#6321] / [Intel XE#8355]) -> [PASS][156]
[155]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-3/igt@xe_evict@evict-mixed-many-threads-small.html
[156]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@xe_evict@evict-mixed-many-threads-small.html
* igt@xe_exec_system_allocator@partial-atomic-middle-remap-cpu-fault:
- shard-bmg: [FAIL][157] ([Intel XE#9115]) -> [PASS][158]
[157]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-2/igt@xe_exec_system_allocator@partial-atomic-middle-remap-cpu-fault.html
[158]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-6/igt@xe_exec_system_allocator@partial-atomic-middle-remap-cpu-fault.html
* igt@xe_exec_system_allocator@threads-many-execqueues-free-nomemset:
- shard-bmg: [ABORT][159] ([Intel XE#6652]) -> [PASS][160]
[159]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-5/igt@xe_exec_system_allocator@threads-many-execqueues-free-nomemset.html
[160]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-2/igt@xe_exec_system_allocator@threads-many-execqueues-free-nomemset.html
* igt@xe_fault_injection@inject-fault-probe-function-xe_wopcm_init:
- shard-bmg: [ABORT][161] ([Intel XE#8007]) -> [PASS][162]
[161]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-1/igt@xe_fault_injection@inject-fault-probe-function-xe_wopcm_init.html
[162]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-7/igt@xe_fault_injection@inject-fault-probe-function-xe_wopcm_init.html
* igt@xe_pm@s4-basic:
- shard-lnl: [ABORT][163] ([Intel XE#9362]) -> [PASS][164] +1 other test pass
[163]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-lnl-5/igt@xe_pm@s4-basic.html
[164]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-lnl-3/igt@xe_pm@s4-basic.html
#### Warnings ####
* igt@kms_async_flips@async-flip-suspend-resume:
- shard-ptl: [ABORT][165] ([Intel XE#9362]) -> [ABORT][166] ([Intel XE#9387])
[165]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-6/igt@kms_async_flips@async-flip-suspend-resume.html
[166]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-1/igt@kms_async_flips@async-flip-suspend-resume.html
* igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs:
- shard-bmg: [INCOMPLETE][167] ([Intel XE#7084] / [Intel XE#8150]) -> [ABORT][168] ([Intel XE#9362])
[167]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-7/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs.html
[168]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-9/igt@kms_ccs@crc-primary-suspend-4-tiled-bmg-ccs.html
* igt@kms_cursor_crc@cursor-suspend:
- shard-ptl: [ABORT][169] ([Intel XE#9362]) -> [ABORT][170] ([Intel XE#9375])
[169]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-ptl-6/igt@kms_cursor_crc@cursor-suspend.html
[170]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-ptl-8/igt@kms_cursor_crc@cursor-suspend.html
* igt@kms_hdr@brightness-with-hdr:
- shard-bmg: [SKIP][171] ([Intel XE#3544]) -> [SKIP][172] ([Intel XE#3374] / [Intel XE#3544])
[171]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-4/igt@kms_hdr@brightness-with-hdr.html
[172]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-8/igt@kms_hdr@brightness-with-hdr.html
* igt@kms_tiled_display@basic-test-pattern-with-chamelium:
- shard-bmg: [SKIP][173] ([Intel XE#2426] / [Intel XE#5848]) -> [SKIP][174] ([Intel XE#2509] / [Intel XE#7437])
[173]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-8/igt@kms_tiled_display@basic-test-pattern-with-chamelium.html
[174]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-1/igt@kms_tiled_display@basic-test-pattern-with-chamelium.html
* igt@xe_pat@pat-sw-hw-suspend:
- shard-bmg: [ABORT][175] ([Intel XE#9362]) -> [FAIL][176] ([Intel XE#7695])
[175]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5824-999838292166407bbe911a41dee23112c49e9d44/shard-bmg-8/igt@xe_pat@pat-sw-hw-suspend.html
[176]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/shard-bmg-1/igt@xe_pat@pat-sw-hw-suspend.html
{name}: This element is suppressed. This means it is ignored when computing
the status of the difference (SUCCESS, WARNING, or FAILURE).
[Intel XE#1124]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1124
[Intel XE#1178]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1178
[Intel XE#1392]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1392
[Intel XE#1397]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1397
[Intel XE#1406]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1406
[Intel XE#1407]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1407
[Intel XE#1421]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1421
[Intel XE#1424]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1424
[Intel XE#1439]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1439
[Intel XE#1489]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1489
[Intel XE#1727]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1727
[Intel XE#1745]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1745
[Intel XE#2234]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2234
[Intel XE#2252]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2252
[Intel XE#2284]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2284
[Intel XE#2311]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2311
[Intel XE#2313]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2313
[Intel XE#2316]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2316
[Intel XE#2320]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2320
[Intel XE#2321]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2321
[Intel XE#2322]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2322
[Intel XE#2327]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2327
[Intel XE#2328]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2328
[Intel XE#2373]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2373
[Intel XE#2387]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2387
[Intel XE#2426]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2426
[Intel XE#2509]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2509
[Intel XE#2763]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2763
[Intel XE#2850]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2850
[Intel XE#2887]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2887
[Intel XE#2893]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2893
[Intel XE#301]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/301
[Intel XE#306]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/306
[Intel XE#3149]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3149
[Intel XE#323]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/323
[Intel XE#3304]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3304
[Intel XE#3374]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3374
[Intel XE#3432]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3432
[Intel XE#3544]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3544
[Intel XE#3718]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3718
[Intel XE#373]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/373
[Intel XE#4141]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/4141
[Intel XE#4156]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/4156
[Intel XE#4227]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/4227
[Intel XE#4298]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/4298
[Intel XE#4733]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/4733
[Intel XE#5575]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5575
[Intel XE#5764]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5764
[Intel XE#5800]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5800
[Intel XE#5808]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5808
[Intel XE#5810]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5810
[Intel XE#5812]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5812
[Intel XE#5822]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5822
[Intel XE#5824]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5824
[Intel XE#584]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/584
[Intel XE#5848]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5848
[Intel XE#5873]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5873
[Intel XE#5900]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5900
[Intel XE#5907]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/5907
[Intel XE#6035]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6035
[Intel XE#6078]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6078
[Intel XE#6312]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6312
[Intel XE#6321]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6321
[Intel XE#651]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/651
[Intel XE#6540]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6540
[Intel XE#656]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/656
[Intel XE#6652]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6652
[Intel XE#6819]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6819
[Intel XE#6867]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6867
[Intel XE#688]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/688
[Intel XE#6886]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6886
[Intel XE#6927]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6927
[Intel XE#6964]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6964
[Intel XE#6969]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6969
[Intel XE#7006]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7006
[Intel XE#7061]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7061
[Intel XE#7084]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7084
[Intel XE#7178]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7178
[Intel XE#7283]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7283
[Intel XE#7304]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7304
[Intel XE#7351]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7351
[Intel XE#7355]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7355
[Intel XE#7356]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7356
[Intel XE#7358]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7358
[Intel XE#7367]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7367
[Intel XE#7369]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7369
[Intel XE#7370]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7370
[Intel XE#7372]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7372
[Intel XE#7374]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7374
[Intel XE#7385]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7385
[Intel XE#7397]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7397
[Intel XE#7402]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7402
[Intel XE#7417]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7417
[Intel XE#7425]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7425
[Intel XE#7429]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7429
[Intel XE#7437]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7437
[Intel XE#7448]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7448
[Intel XE#7482]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7482
[Intel XE#7484]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7484
[Intel XE#7590]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7590
[Intel XE#7599]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7599
[Intel XE#7642]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7642
[Intel XE#7679]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7679
[Intel XE#7695]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7695
[Intel XE#7724]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7724
[Intel XE#7772]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7772
[Intel XE#7793]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7793
[Intel XE#7865]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7865
[Intel XE#7905]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7905
[Intel XE#8007]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8007
[Intel XE#8150]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8150
[Intel XE#8265]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8265
[Intel XE#8339]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8339
[Intel XE#8355]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8355
[Intel XE#836]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/836
[Intel XE#8364]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8364
[Intel XE#8369]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8369
[Intel XE#8374]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8374
[Intel XE#8378]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8378
[Intel XE#8854]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8854
[Intel XE#8963]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8963
[Intel XE#9096]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9096
[Intel XE#9115]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9115
[Intel XE#9144]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9144
[Intel XE#9333]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9333
[Intel XE#9362]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9362
[Intel XE#9367]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9367
[Intel XE#9375]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9375
[Intel XE#9387]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9387
[Intel XE#9396]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9396
[Intel XE#9399]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/9399
[Intel XE#944]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/944
Build changes
-------------
* Linux: xe-5824-999838292166407bbe911a41dee23112c49e9d44 -> xe-pw-149888v9
IGT_9112: 8d3284b5ae16ca1aef353dfaffb6c0896ceacc19 @ https://gitlab.freedesktop.org/drm/igt-gpu-tools.git
xe-5824-999838292166407bbe911a41dee23112c49e9d44: 999838292166407bbe911a41dee23112c49e9d44
xe-pw-149888v9: 149888v9
== Logs ==
For more details see: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-149888v9/index.html
[-- Attachment #2: Type: text/html, Size: 62403 bytes --]
^ permalink raw reply [flat|nested] 54+ messages in thread* Re: [PATCH v7 00/24] CPU binds and ULLS on migration queue
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
` (27 preceding siblings ...)
2026-09-25 15:05 ` ✗ Xe.CI.FULL: failure " Patchwork
@ 2026-09-25 18:08 ` Maarten Lankhorst
28 siblings, 0 replies; 54+ messages in thread
From: Maarten Lankhorst @ 2026-09-25 18:08 UTC (permalink / raw)
To: Matthew Brost, intel-xe
I've run with a previous version of this series in my local tree for quite some time.
Hoping this finally gets merged!
Reviewed-by: Maarten Lankhorst <dev@lankhorst.se> for the whole series.
On 9/25/26 06:52, Matthew Brost wrote:
> We now have data demonstrating the need for CPU binds and ULLS on the
> migration queue, based on results generated from [1].
>
> On BMG, measurements show that when the GPU is continuously processing
> faults, copy jobs run approximately 30–40µs faster (depending on the
> test case) with ULLS compared to traditional GuC submission with SLPC
> enabled on the migration queue. Startup from a cold GPU shows an even
> larger speedup. Given the critical nature of fault performance, ULLS
> appears to be a worthwhile feature.
>
> In addition to driver telemetry, UMD compute benchmarks consistently
> show multiple GB/s improvement in pagefault benchmarks with ULLS enabled.
>
> ULLS will consume more power (not yet measured) due to a continuously
> running batch on the paging engine. However, compute UMDs already do
> this on engines exposed to users, so this seems like a worthwhile
> tradeoff. To mitigate power concerns, ULLS will exit after a period of
> time in which no faults have been processed.
>
> CPU binds are required for ULLS to function, as the migration queue
> needs exclusive access to the paging hardware engine. Thus, CPU binds
> are included here.
>
> Beyond being a requirement for ULLS, CPU binds should also reduce
> VM-bind latency, provide clearer multi-tile and TLB-invalidation
> layering, reduce pressure on GuC during fault storms as it is bypassed,
> and decouple kernel binds from unrelated copy/clear jobs—especially
> beneficial when faults are serviced in parallel. In a parallel-faulting
> test case, average bind time was reduced by approximately 15µs. In the
> worst case, 2MB copy time (~60–140µs) × (number of pagefault threads −
> 1) of latency would otherwise be added to a single fault. Reducing this
> latency increases overall throughput of the fault handler.
>
> This series can be merged in phases:
>
> Phase 1: CPU binds (patches 1–13)
> Phase 2: CPU-bind components and multi-tile relayers (patches 14–17)
> Phase 3: ULLS on the migration execution queue (patches 18–24)
>
> v2:
> - Use delayed worker to exit ULLS mode in an effort to save on power
> - Various other cleanups
> v3:
> - CPU bind component, multi-tile relayer
> - Split CPU bind patches in many small patches
> v4:
> - Rebase, address feedback, add ULLS doc patch
> v5:
> - Address Sashiko feedback, fix checkpatch issues
> v6:
> - Reworked ULLS to drop KMD MMIO ring tail write, move this to ring
> instruction(s)
> - Several other small cleanups based on Sashiko feedback
> v7:
> - Drop XE_BO_FLAG_PUT_VM_ASYNC in favor PT BOs holding a VM ref
> - configfs to ULLS enable and period
> - Migrate ULLS implementation style cleanups
> - Address various feedback
>
> Matt
>
> [1] https://patchwork.freedesktop.org/series/149811/
>
> Matthew Brost (24):
> drm/xe: reference VM from PT BOs
> drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear
> vfuns
> drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper
> drm/xe: Decouple exec queue idle check from LRC
> drm/xe: Add job count to GuC exec queue snapshot
> drm/xe: Update xe_bo_put_deferred arguments to include writeback flag
> drm/xe: Update scheduler job layer to support PT jobs
> drm/xe: Add helpers to access PT ops
> drm/xe: Add struct xe_pt_job_ops
> drm/xe: Update GuC submission backend to run PT jobs
> drm/xe: Store level in struct xe_vm_pgtable_update
> drm/xe: Don't use migrate exec queue for page fault binds
> drm/xe: Enable CPU binds for jobs
> drm/xe: Remove unused arguments from xe_migrate_pt_update_ops
> drm/xe: Make bind queues operate cross-tile
> drm/xe: Add CPU bind layer
> drm/xe: Add device flag to enable PT mirroring across tiles
> drm/xe: Add ULLS support to LRC
> drm/xe: Add ULLS migration job support to migration layer
> drm/xe: Add ULLS migration job support to ring ops
> drm/xe: Add ULLS migration job support to GuC submission
> drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch
> drm/xe: add migrate ULLS period configfs attribute
> drm/xe: Document ULLS for migration jobs
>
> Documentation/gpu/xe/xe_migrate.rst | 3 +
> drivers/gpu/drm/xe/Makefile | 1 +
> drivers/gpu/drm/xe/xe_bo.c | 4 +-
> drivers/gpu/drm/xe/xe_bo.h | 10 +-
> drivers/gpu/drm/xe/xe_bo_types.h | 2 -
> drivers/gpu/drm/xe/xe_configfs.c | 67 ++
> drivers/gpu/drm/xe/xe_configfs.h | 2 +
> drivers/gpu/drm/xe/xe_cpu_bind.c | 296 +++++++++
> drivers/gpu/drm/xe/xe_cpu_bind.h | 111 ++++
> drivers/gpu/drm/xe/xe_device.c | 5 +
> drivers/gpu/drm/xe/xe_device_types.h | 6 +
> drivers/gpu/drm/xe/xe_drm_client.c | 2 +-
> drivers/gpu/drm/xe/xe_exec_queue.c | 163 ++---
> drivers/gpu/drm/xe/xe_exec_queue.h | 16 +-
> drivers/gpu/drm/xe/xe_exec_queue_types.h | 20 +-
> drivers/gpu/drm/xe/xe_guc_submit.c | 76 ++-
> drivers/gpu/drm/xe/xe_guc_submit_types.h | 2 +
> drivers/gpu/drm/xe/xe_lrc.c | 73 +++
> drivers/gpu/drm/xe/xe_lrc.h | 4 +
> drivers/gpu/drm/xe/xe_lrc_types.h | 4 +
> drivers/gpu/drm/xe/xe_migrate.c | 769 ++++++++++------------
> drivers/gpu/drm/xe/xe_migrate.h | 95 +--
> drivers/gpu/drm/xe/xe_pagefault.c | 3 +
> drivers/gpu/drm/xe/xe_pci.c | 2 +
> drivers/gpu/drm/xe/xe_pci_types.h | 3 +-
> drivers/gpu/drm/xe/xe_pt.c | 781 ++++++++++++++---------
> drivers/gpu/drm/xe/xe_pt.h | 12 +-
> drivers/gpu/drm/xe/xe_pt_types.h | 49 +-
> drivers/gpu/drm/xe/xe_ring_ops.c | 75 +++
> drivers/gpu/drm/xe/xe_ring_ops_types.h | 24 +
> drivers/gpu/drm/xe/xe_sched_job.c | 103 ++-
> drivers/gpu/drm/xe/xe_sched_job.h | 56 ++
> drivers/gpu/drm/xe/xe_sched_job_types.h | 49 +-
> drivers/gpu/drm/xe/xe_sync.c | 20 +-
> drivers/gpu/drm/xe/xe_tlb_inval_job.c | 28 +-
> drivers/gpu/drm/xe/xe_tlb_inval_job.h | 4 +-
> drivers/gpu/drm/xe/xe_trace.h | 2 +-
> drivers/gpu/drm/xe/xe_vm.c | 242 +++----
> drivers/gpu/drm/xe/xe_vm.h | 3 +
> drivers/gpu/drm/xe/xe_vm_types.h | 12 +-
> 40 files changed, 1988 insertions(+), 1211 deletions(-)
> create mode 100644 drivers/gpu/drm/xe/xe_cpu_bind.c
> create mode 100644 drivers/gpu/drm/xe/xe_cpu_bind.h
>
^ permalink raw reply [flat|nested] 54+ messages in thread