From: Matthew Brost <matthew.brost@intel.com>
To: <sashiko-reviews@lists.linux.dev>
Cc: <intel-xe@lists.freedesktop.org>
Subject: Re: [PATCH v4 10/25] drm/xe: Update GuC submission backend to run PT jobs
Date: Thu, 3 Sep 2026 17:57:53 -0700 [thread overview]
Message-ID: <apoXkWK7ybUlsKIx@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260904003650.4686B1F000E9@smtp.kernel.org>
On Fri, Sep 04, 2026 at 12:36:49AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [High] PT jobs on wedged, killed, or banned queues silently report success to the DRM scheduler, masking failures.
> - [High] Memory leak of xe_pt_job_ops and dma_fence references if a PT job is aborted before execution.
> --
>
> commit 86f7996687fb641831b3ca6cb60d0fed641a6978
> Author: Matthew Brost <matthew.brost@intel.com>
>
> drm/xe: Update GuC submission backend to run PT jobs
>
> This commit teaches the GuC submission backend to execute PT (Page Table) jobs.
> PT jobs bypass GPU execution for the final step of a bind job, instead using
> the CPU to program the required page tables.
>
> Link: https://patch.msgid.link/20260228013501.106680-11-matthew.brost@intel.com
>
> > diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> > --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> > +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
> > @@ -1261,14 +1277,25 @@ guc_exec_queue_run_job(struct drm_sched_job *drm_job)
>
> [ ... ]
>
> > if (!killed_or_banned_or_wedged && !xe_sched_job_is_error(job)) {
>
> [ ... ]
>
> > + if (is_pt_job(job)) {
> > + xe_gt_assert(guc_to_gt(guc), !exec_queue_registered(q));
> > + run_pt_job(job);
> > + } else {
> > + if (!exec_queue_registered(q))
> > + register_exec_queue(q, GUC_CONTEXT_NORMAL);
> > + if (!job->restore_replay)
> > + q->ring_ops->emit_job(job);
> > + submit_exec_queue(q, job);
> > + }
> > job->restore_replay = false;
> > }
> >
> > + if (is_pt_job(job)) {
> > + xe_pt_job_ops_put(job->pt_update[0].pt_job_ops);
> > + dma_fence_put(job->fence); /* Drop ref from xe_sched_job_arm */
> > + return NULL;
> > + }
>
> [Severity: High]
> If the queue is wedged, killed, or banned, the execution block above is
> skipped, but this block unconditionally returns NULL for PT jobs. Can this
> cause PT jobs on failed queues to silently report success to the DRM
> scheduler and mask failures?
>
The intent to is drop the job on sliently in this case, so code is correct.
> [Severity: High]
> This handles dropping the extra dma_fence and xe_pt_job_ops references during
> normal execution, but what happens if the DRM scheduler aborts the job before
> it executes (for example, if the queue is destroyed)?
>
> Looking at xe_sched_job_destroy() in drivers/gpu/drm/xe/xe_sched_job.c:
>
> > void xe_sched_job_destroy(struct kref *ref)
> > {
> > struct xe_sched_job *job =
> > container_of(ref, struct xe_sched_job, refcount);
> > ...
> > xe_sched_job_free_fences(job);
> > dma_fence_put(job->fence);
> > drm_sched_job_cleanup(&job->drm);
> > job_free(job);
> > ...
> > }
>
> It seems the destructor fails to call xe_pt_job_ops_put() for PT jobs and
> only drops the primary fence reference. Does this code leak the xe_pt_job_ops
> structure and the extra dma_fence reference taken during job arming if
> execution is aborted?
Every job put into the scheduler reaches run_job, so a non-issue.
Matt
>
> > run_job_out:
> >
> > return job->fence;
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260903235842.3401722-1-matthew.brost@intel.com?part=10
next prev parent reply other threads:[~2026-09-04 0:58 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 23:58 [PATCH v4 00/25] CPU binds and ULLS on migration queue Matthew Brost
2026-09-03 23:58 ` [PATCH v4 01/25] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
2026-09-03 23:58 ` [PATCH v4 02/25] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
2026-09-04 0:15 ` sashiko-bot
2026-09-03 23:58 ` [PATCH v4 03/25] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
2026-09-03 23:58 ` [PATCH v4 04/25] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
2026-09-03 23:58 ` [PATCH v4 05/25] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
2026-09-03 23:58 ` [PATCH v4 06/25] drm/xe: Add XE_BO_FLAG_PUT_VM_ASYNC Matthew Brost
2026-09-04 0:18 ` sashiko-bot
2026-09-04 0:41 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 07/25] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
2026-09-04 0:25 ` sashiko-bot
2026-09-03 23:58 ` [PATCH v4 08/25] drm/xe: Add helpers to access PT ops Matthew Brost
2026-09-03 23:58 ` [PATCH v4 09/25] drm/xe: Add struct xe_pt_job_ops Matthew Brost
2026-09-03 23:58 ` [PATCH v4 10/25] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
2026-09-04 0:36 ` sashiko-bot
2026-09-04 0:57 ` Matthew Brost [this message]
2026-09-03 23:58 ` [PATCH v4 11/25] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
2026-09-04 0:19 ` sashiko-bot
2026-09-03 23:58 ` [PATCH v4 12/25] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
2026-09-03 23:58 ` [PATCH v4 13/25] drm/xe: Enable CPU binds for jobs Matthew Brost
2026-09-04 0:31 ` sashiko-bot
2026-09-04 1:04 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 14/25] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
2026-09-03 23:58 ` [PATCH v4 15/25] drm/xe: Make bind queues operate cross-tile Matthew Brost
2026-09-03 23:58 ` [PATCH v4 16/25] drm/xe: Add CPU bind layer Matthew Brost
2026-09-04 0:31 ` sashiko-bot
2026-09-04 1:18 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 17/25] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
2026-09-04 0:29 ` sashiko-bot
2026-09-04 1:33 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 18/25] drm/xe: Add xe_hw_engine_write_ring_tail Matthew Brost
2026-09-03 23:58 ` [PATCH v4 19/25] drm/xe: Add ULLS support to LRC Matthew Brost
2026-09-03 23:58 ` [PATCH v4 20/25] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
2026-09-04 0:27 ` sashiko-bot
2026-09-04 1:35 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 21/25] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
2026-09-03 23:58 ` [PATCH v4 22/25] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
2026-09-04 0:38 ` sashiko-bot
2026-09-04 1:41 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 23/25] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
2026-09-04 0:28 ` sashiko-bot
2026-09-04 1:32 ` Matthew Brost
2026-09-03 23:58 ` [PATCH v4 24/25] drm/xe: Add modparam to enable / disable ULLS on migrate queue Matthew Brost
2026-09-03 23:58 ` [PATCH v4 25/25] drm/xe: Document ULLS for migration jobs Matthew Brost
2026-09-04 0:47 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev6) Patchwork
2026-09-04 0:49 ` ✓ CI.KUnit: success " Patchwork
2026-09-04 1:33 ` ✓ Xe.CI.BAT: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apoXkWK7ybUlsKIx@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox