Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Matthew Brost" <matthew.brost@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs
Date: Fri, 25 Sep 2026 06:03:25 +0000	[thread overview]
Message-ID: <20260925060326.27F9A1F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260925045320.1325860-14-matthew.brost@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] drm/xe: memory leak of pt_job_ops and dma_fence when PT jobs are canceled
- [High] drm/xe: GPU hang and memory leak due to missing CPU bind support in Execlist backend
--

commit 6471c15e2149c2517458e891e7dceef8c5573f97
Author: Matthew Brost <matthew.brost@intel.com>

drm/xe: Enable CPU binds for jobs

This commit updates the migration layer to formulate page table jobs that
issue CPU binds in the submission backend. All code related to GPU-based
binding is removed to lower latency and avoid PTE migration limits.

> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
[ ... ]
> @@ -1937,162 +1866,18 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
>  {
>  	const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
>  	struct xe_tile *tile = m->tile;
>  	struct xe_sched_job *job;
>  	struct dma_fence *fence;
>  	bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
> +	int err;
>  
> -	job = xe_bb_create_migration_job(pt_update_ops->q, bb,
> -					 xe_migrate_batch_base(m, usm),
> -					 update_idx);
> +	job = xe_sched_job_create(pt_update_ops->q, NULL);

[Severity: High]
Does the Execlist backend properly handle these new PT jobs?

Since xe_sched_job_create() is now called with a NULL batch address, these
jobs rely on backend interception. However, looking at execlist_run_job() in
the Execlist backend:

drivers/gpu/drm/xe/xe_execlist.c:execlist_run_job() {
    ...
    q->ring_ops->emit_job(job);
    xe_execlist_make_active(exl);

    return job->fence;
}

If GuC submission is disabled, it appears the backend receives this job and
directly emits it to the ring without an is_pt_job() check. Could this result
in pushing an uninitialized batch address to the hardware, causing a GPU hang
and leaking the pt_job_ops structures?

>  	if (IS_ERR(job)) {
>  		err = PTR_ERR(job);
> -		goto err_sa;
> +		goto err_out;
>  	}
>  
> -	xe_sched_job_add_migrate_flush(job, MI_INVALIDATE_TLB);
> +	xe_tile_assert(tile, job->is_pt_job);
>  
>  	if (ops->pre_commit) {
>  		pt_update->job = job;
> @@ -2103,6 +1888,12 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
>  	if (is_migrate)
>  		mutex_lock(&m->job_mutex);
>  
> +	job->pt_update[0].vm = pt_update->vops->vm;
> +	job->pt_update[0].tile = tile;
> +	job->pt_update[0].ops = ops;
> +	job->pt_update[0].pt_job_ops =
> +		xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
> +

[Severity: High]
Is there a memory leak if this PT job gets canceled before execution?

We acquire a reference to pt_job_ops here, and xe_sched_job_arm() takes an
extra dma_fence reference. If the job is canceled (for example, from an
execution queue kill or GPU reset), the DRM scheduler calls the free_job
callback which ultimately calls xe_sched_job_destroy():

drivers/gpu/drm/xe/xe_sched_job.c:xe_sched_job_destroy() {
    ...
    xe_sched_job_free_fences(job);
    dma_fence_put(job->fence);
    drm_sched_job_cleanup(&job->drm);
    job_free(job);
}

Since xe_sched_job_destroy() doesn't check job->is_pt_job, it seems we never
drop the pt_job_ops or the extra dma_fence reference. This would permanently
leak the deferred actions and their associated page table BO references. Does
the job destruction path need updating to clean up canceled PT jobs?

>  	xe_sched_job_arm(job);
>  	fence = dma_fence_get(&job->drm.s_fence->finished);
>  	xe_sched_job_push(job);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=13

  reply	other threads:[~2026-09-25  6:03 UTC|newest]

Thread overview: 54+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25  4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
2026-09-25  4:52 ` [PATCH v7 01/24] drm/xe: reference VM from PT BOs Matthew Brost
2026-09-25 12:00   ` Francois Dugast
2026-09-25 16:10     ` Matthew Brost
2026-09-25  4:52 ` [PATCH v7 02/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
2026-09-25  4:52 ` [PATCH v7 03/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
2026-09-25  4:53 ` [PATCH v7 04/24] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
2026-09-25  4:53 ` [PATCH v7 05/24] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
2026-09-25  4:53 ` [PATCH v7 06/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
2026-09-25  4:53 ` [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
2026-09-25 11:23   ` Francois Dugast
2026-09-25  4:53 ` [PATCH v7 08/24] drm/xe: Add helpers to access PT ops Matthew Brost
2026-09-25  4:53 ` [PATCH v7 09/24] drm/xe: Add struct xe_pt_job_ops Matthew Brost
2026-09-25  4:53 ` [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
2026-09-25  5:39   ` sashiko-bot
2026-09-25  5:58     ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 11/24] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
2026-09-25  4:53 ` [PATCH v7 12/24] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
2026-09-25  4:53 ` [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
2026-09-25  6:03   ` sashiko-bot [this message]
2026-09-25  6:54     ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
2026-09-25  4:53 ` [PATCH v7 15/24] drm/xe: Make bind queues operate cross-tile Matthew Brost
2026-09-25  4:53 ` [PATCH v7 16/24] drm/xe: Add CPU bind layer Matthew Brost
2026-09-25  4:53 ` [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
2026-09-25  6:25   ` sashiko-bot
2026-09-25  7:21     ` Matthew Brost
2026-09-25  9:51   ` Francois Dugast
2026-09-25  4:53 ` [PATCH v7 18/24] drm/xe: Add ULLS support to LRC Matthew Brost
2026-09-25  4:53 ` [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
2026-09-25  6:34   ` sashiko-bot
2026-09-25  7:17     ` Matthew Brost
2026-09-25 20:10       ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
2026-09-25  6:38   ` sashiko-bot
2026-09-25  7:02     ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
2026-09-25  6:48   ` sashiko-bot
2026-09-25  7:08     ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
2026-09-25 17:49   ` Maarten Lankhorst
2026-09-25 18:17     ` Matthew Brost
2026-09-25 18:26       ` Maarten Lankhorst
2026-09-25 19:38         ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute Matthew Brost
2026-09-25  6:51   ` sashiko-bot
2026-09-25  7:08     ` Matthew Brost
2026-09-25  4:53 ` [PATCH v7 24/24] drm/xe: Document ULLS for migration jobs Matthew Brost
2026-09-25  5:02 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev9) Patchwork
2026-09-25  5:04 ` ✓ CI.KUnit: success " Patchwork
2026-09-25  5:47 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-25 15:05 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-25 16:01   ` Matthew Brost
2026-09-25 18:08 ` [PATCH v7 00/24] CPU binds and ULLS on migration queue Maarten Lankhorst

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260925060326.27F9A1F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox