Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Matthew Brost" <matthew.brost@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v8 10/24] drm/xe: Update GuC submission backend to run PT jobs
Date: Fri, 25 Sep 2026 20:38:51 +0000	[thread overview]
Message-ID: <20260925203851.D655B1F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260925201526.1355279-11-matthew.brost@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] drm/xe: Memory leak of pt_job_ops and fence refs on aborted PT jobs
- [High] drm/xe: guc_exec_queue_run_job() drops error status for aborted PT jobs
--

commit 9cf187bc8ca7b690be92bb4415979a7f9395643d
Author: Matthew Brost <matthew.brost@intel.com>
Date:   Fri Sep 25 20:15:12 2026 +0000

drm/xe: Update GuC submission backend to run PT jobs

PT jobs bypass GPU execution for the final step of a bind job, using the
CPU to program the required page tables. Teach the GuC submission backend
how to execute these jobs.

> diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
[ ... ]
>  static struct dma_fence *
>  guc_exec_queue_run_job(struct drm_sched_job *drm_job)
>  {
>  	struct xe_sched_job *job = to_xe_sched_job(drm_job);
>  	struct xe_exec_queue *q = job->q;
>  	struct xe_guc *guc = exec_queue_to_guc(q);
> -	bool killed_or_banned_or_wedged =
> -		exec_queue_killed_or_banned_or_wedged(q);
> +	bool killed_or_banned_or_wedged_or_error  =
> +		exec_queue_killed_or_banned_or_wedged(q) ||
> +		xe_sched_job_is_error(job);
>  
>  	xe_gt_assert(guc_to_gt(guc), !(exec_queue_destroyed(q) || exec_queue_pending_disable(q)) ||
>  		     exec_queue_banned(q) || exec_queue_suspended(q));
>  
>  	trace_xe_sched_job_run(job);
>  
> -	if (!killed_or_banned_or_wedged && !xe_sched_job_is_error(job)) {
> +	if (is_pt_job(job)) {
> +		xe_gt_assert(guc_to_gt(guc), !exec_queue_registered(q));
> +		run_pt_job(job, killed_or_banned_or_wedged_or_error);
> +		xe_pt_job_ops_put(job->pt_update[0].pt_job_ops);
> +		dma_fence_put(job->fence);	/* Drop ref from xe_sched_job_arm */

[Severity: High]
If a PT job is aborted by the DRM scheduler before guc_exec_queue_run_job()
is called, do these references get properly released?

When jobs are aborted (e.g., due to failing dependencies or context
teardown), the cleanup path typically goes through xe_sched_job_destroy().

Since these references are dropped exclusively on this execution path, it
seems they might permanently leak xe_pt_job_ops and dma_fence_stub if the
job never makes it here.

> +
> +		return NULL;

[Severity: High]
Could unconditionally returning NULL here hide error states from the DRM
scheduler?

When killed_or_banned_or_wedged_or_error is true, run_pt_job() is passed
true for force_clear, which correctly clears the page tables. 

However, returning NULL tells the DRM scheduler that the job completed
successfully without error.

This false success signal might cause dependent jobs on other queues to
execute on unbound memory and trigger GPU page faults.

> +	} else if (!killed_or_banned_or_wedged_or_error) {
>  		if (xe_exec_queue_is_multi_queue_secondary(q)) {

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260925201526.1355279-1-matthew.brost@intel.com?part=10

  reply	other threads:[~2026-09-25 20:38 UTC|newest]

Thread overview: 36+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25 20:15 [PATCH v8 00/24] CPU binds and ULLS on migration queue Matthew Brost
2026-09-25 20:15 ` [PATCH v8 01/24] drm/xe: reference VM from PT BOs Matthew Brost
2026-09-25 20:15 ` [PATCH v8 02/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
2026-09-25 20:15 ` [PATCH v8 03/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
2026-09-25 20:15 ` [PATCH v8 04/24] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
2026-09-25 20:15 ` [PATCH v8 05/24] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
2026-09-25 20:15 ` [PATCH v8 06/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
2026-09-25 20:15 ` [PATCH v8 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
2026-09-25 20:15 ` [PATCH v8 08/24] drm/xe: Add helpers to access PT ops Matthew Brost
2026-09-25 20:15 ` [PATCH v8 09/24] drm/xe: Add struct xe_pt_job_ops Matthew Brost
2026-09-25 20:15 ` [PATCH v8 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
2026-09-25 20:38   ` sashiko-bot [this message]
2026-09-25 23:20   ` Ghimiray, Himal Prasad
2026-09-25 20:15 ` [PATCH v8 11/24] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
2026-09-25 20:15 ` [PATCH v8 12/24] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
2026-09-25 20:15 ` [PATCH v8 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
2026-09-25 20:39   ` sashiko-bot
2026-09-25 20:15 ` [PATCH v8 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
2026-09-25 20:15 ` [PATCH v8 15/24] drm/xe: Make bind queues operate cross-tile Matthew Brost
2026-09-25 20:15 ` [PATCH v8 16/24] drm/xe: Add CPU bind layer Matthew Brost
2026-09-25 20:15 ` [PATCH v8 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
2026-09-25 20:46   ` sashiko-bot
2026-09-25 20:15 ` [PATCH v8 18/24] drm/xe: Add ULLS support to LRC Matthew Brost
2026-09-25 20:15 ` [PATCH v8 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
2026-09-25 20:41   ` sashiko-bot
2026-09-25 20:15 ` [PATCH v8 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
2026-09-25 20:48   ` sashiko-bot
2026-09-25 20:15 ` [PATCH v8 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
2026-09-25 20:45   ` sashiko-bot
2026-09-25 20:15 ` [PATCH v8 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
2026-09-25 20:15 ` [PATCH v8 23/24] drm/xe: add migrate ULLS period configfs attribute Matthew Brost
2026-09-25 20:15 ` [PATCH v8 24/24] drm/xe: Document ULLS for migration jobs Matthew Brost
2026-09-25 21:02 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev10) Patchwork
2026-09-25 21:04 ` ✓ CI.KUnit: success " Patchwork
2026-09-25 22:12 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-26  7:40 ` ✗ Xe.CI.FULL: failure " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260925203851.D655B1F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox