From: Matthew Brost <matthew.brost@intel.com>
To: <sashiko-reviews@lists.linux.dev>
Cc: <intel-xe@lists.freedesktop.org>
Subject: Re: [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs
Date: Thu, 24 Sep 2026 23:54:28 -0700 [thread overview]
Message-ID: <arYapDIMe2PAW4qF@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260925060326.27F9A1F000FF@smtp.kernel.org>
On Fri, Sep 25, 2026 at 06:03:25AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [High] drm/xe: memory leak of pt_job_ops and dma_fence when PT jobs are canceled
> - [High] drm/xe: GPU hang and memory leak due to missing CPU bind support in Execlist backend
> --
>
> commit 6471c15e2149c2517458e891e7dceef8c5573f97
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Enable CPU binds for jobs
>
> This commit updates the migration layer to formulate page table jobs that
> issue CPU binds in the submission backend. All code related to GPU-based
> binding is removed to lower latency and avoid PTE migration limits.
>
> > diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> > --- a/drivers/gpu/drm/xe/xe_migrate.c
> > +++ b/drivers/gpu/drm/xe/xe_migrate.c
> [ ... ]
> > @@ -1937,162 +1866,18 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
> > {
> > const struct xe_migrate_pt_update_ops *ops = pt_update->ops;
> > struct xe_tile *tile = m->tile;
> > struct xe_sched_job *job;
> > struct dma_fence *fence;
> > bool is_migrate = is_migrate_queue(m, pt_update_ops->q);
> > + int err;
> >
> > - job = xe_bb_create_migration_job(pt_update_ops->q, bb,
> > - xe_migrate_batch_base(m, usm),
> > - update_idx);
> > + job = xe_sched_job_create(pt_update_ops->q, NULL);
>
> [Severity: High]
> Does the Execlist backend properly handle these new PT jobs?
>
Execlist backend is dead. Non-issue.
> Since xe_sched_job_create() is now called with a NULL batch address, these
> jobs rely on backend interception. However, looking at execlist_run_job() in
> the Execlist backend:
>
> drivers/gpu/drm/xe/xe_execlist.c:execlist_run_job() {
> ...
> q->ring_ops->emit_job(job);
> xe_execlist_make_active(exl);
>
> return job->fence;
> }
>
> If GuC submission is disabled, it appears the backend receives this job and
> directly emits it to the ring without an is_pt_job() check. Could this result
> in pushing an uninitialized batch address to the hardware, causing a GPU hang
> and leaking the pt_job_ops structures?
>
> > if (IS_ERR(job)) {
> > err = PTR_ERR(job);
> > - goto err_sa;
> > + goto err_out;
> > }
> >
> > - xe_sched_job_add_migrate_flush(job, MI_INVALIDATE_TLB);
> > + xe_tile_assert(tile, job->is_pt_job);
> >
> > if (ops->pre_commit) {
> > pt_update->job = job;
> > @@ -2103,6 +1888,12 @@ __xe_migrate_update_pgtables(struct xe_migrate *m,
> > if (is_migrate)
> > mutex_lock(&m->job_mutex);
> >
> > + job->pt_update[0].vm = pt_update->vops->vm;
> > + job->pt_update[0].tile = tile;
> > + job->pt_update[0].ops = ops;
> > + job->pt_update[0].pt_job_ops =
> > + xe_pt_job_ops_get(pt_update_ops->pt_job_ops);
> > +
>
> [Severity: High]
> Is there a memory leak if this PT job gets canceled before execution?
>
Again non-issue - _arm() always equals run_job() call. Non-issue.
Matt
> We acquire a reference to pt_job_ops here, and xe_sched_job_arm() takes an
> extra dma_fence reference. If the job is canceled (for example, from an
> execution queue kill or GPU reset), the DRM scheduler calls the free_job
> callback which ultimately calls xe_sched_job_destroy():
>
> drivers/gpu/drm/xe/xe_sched_job.c:xe_sched_job_destroy() {
> ...
> xe_sched_job_free_fences(job);
> dma_fence_put(job->fence);
> drm_sched_job_cleanup(&job->drm);
> job_free(job);
> }
>
> Since xe_sched_job_destroy() doesn't check job->is_pt_job, it seems we never
> drop the pt_job_ops or the extra dma_fence reference. This would permanently
> leak the deferred actions and their associated page table BO references. Does
> the job destruction path need updating to clean up canceled PT jobs?
>
> > xe_sched_job_arm(job);
> > fence = dma_fence_get(&job->drm.s_fence->finished);
> > xe_sched_job_push(job);
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260925045320.1325860-1-matthew.brost@intel.com?part=13
next prev parent reply other threads:[~2026-09-25 6:55 UTC|newest]
Thread overview: 54+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-25 4:52 [PATCH v7 00/24] CPU binds and ULLS on migration queue Matthew Brost
2026-09-25 4:52 ` [PATCH v7 01/24] drm/xe: reference VM from PT BOs Matthew Brost
2026-09-25 12:00 ` Francois Dugast
2026-09-25 16:10 ` Matthew Brost
2026-09-25 4:52 ` [PATCH v7 02/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
2026-09-25 4:52 ` [PATCH v7 03/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
2026-09-25 4:53 ` [PATCH v7 04/24] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
2026-09-25 4:53 ` [PATCH v7 05/24] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
2026-09-25 4:53 ` [PATCH v7 06/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
2026-09-25 4:53 ` [PATCH v7 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
2026-09-25 11:23 ` Francois Dugast
2026-09-25 4:53 ` [PATCH v7 08/24] drm/xe: Add helpers to access PT ops Matthew Brost
2026-09-25 4:53 ` [PATCH v7 09/24] drm/xe: Add struct xe_pt_job_ops Matthew Brost
2026-09-25 4:53 ` [PATCH v7 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
2026-09-25 5:39 ` sashiko-bot
2026-09-25 5:58 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 11/24] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
2026-09-25 4:53 ` [PATCH v7 12/24] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
2026-09-25 4:53 ` [PATCH v7 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
2026-09-25 6:03 ` sashiko-bot
2026-09-25 6:54 ` Matthew Brost [this message]
2026-09-25 4:53 ` [PATCH v7 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
2026-09-25 4:53 ` [PATCH v7 15/24] drm/xe: Make bind queues operate cross-tile Matthew Brost
2026-09-25 4:53 ` [PATCH v7 16/24] drm/xe: Add CPU bind layer Matthew Brost
2026-09-25 4:53 ` [PATCH v7 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
2026-09-25 6:25 ` sashiko-bot
2026-09-25 7:21 ` Matthew Brost
2026-09-25 9:51 ` Francois Dugast
2026-09-25 4:53 ` [PATCH v7 18/24] drm/xe: Add ULLS support to LRC Matthew Brost
2026-09-25 4:53 ` [PATCH v7 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
2026-09-25 6:34 ` sashiko-bot
2026-09-25 7:17 ` Matthew Brost
2026-09-25 20:10 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
2026-09-25 6:38 ` sashiko-bot
2026-09-25 7:02 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
2026-09-25 6:48 ` sashiko-bot
2026-09-25 7:08 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
2026-09-25 17:49 ` Maarten Lankhorst
2026-09-25 18:17 ` Matthew Brost
2026-09-25 18:26 ` Maarten Lankhorst
2026-09-25 19:38 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 23/24] drm/xe: add migrate ULLS period configfs attribute Matthew Brost
2026-09-25 6:51 ` sashiko-bot
2026-09-25 7:08 ` Matthew Brost
2026-09-25 4:53 ` [PATCH v7 24/24] drm/xe: Document ULLS for migration jobs Matthew Brost
2026-09-25 5:02 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev9) Patchwork
2026-09-25 5:04 ` ✓ CI.KUnit: success " Patchwork
2026-09-25 5:47 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-25 15:05 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-25 16:01 ` Matthew Brost
2026-09-25 18:08 ` [PATCH v7 00/24] CPU binds and ULLS on migration queue Maarten Lankhorst
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arYapDIMe2PAW4qF@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox