Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Matthew Brost" <matthew.brost@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v6 09/24] drm/xe: Add struct xe_pt_job_ops
Date: Fri, 04 Sep 2026 21:40:02 +0000	[thread overview]
Message-ID: <20260904214003.09F171F00A3D@smtp.kernel.org> (raw)
In-Reply-To: <20260904211613.3934307-10-matthew.brost@intel.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Missing BO refcount re-initialization causes refcount underflow and memory leak during async free.
--

commit 6e95dc31189c14b0582d03e7feb8e249921d75d3
Author: Matthew Brost <matthew.brost@intel.com>
Subject: drm/xe: Add struct xe_pt_job_ops

This commit introduces struct xe_pt_job_ops, a dynamically refcounted object
containing the information required to issue a CPU bind via a job after the
initial bind IOCTL returns.

Link: https://patch.msgid.link/20260228013501.106680-10-matthew.brost@intel.com

> diff --git a/drivers/gpu/drm/xe/xe_pt.c b/drivers/gpu/drm/xe/xe_pt.c
> index e560f167fdf1b..30127ebf1b605 100644
> --- a/drivers/gpu/drm/xe/xe_pt.c
> +++ b/drivers/gpu/drm/xe/xe_pt.c

[ ... ]

> @@ -213,7 +214,19 @@ void xe_pt_destroy(struct xe_pt *pt, u32 flags, struct llist_head *deferred)
>  
>  	XE_WARN_ON(!list_empty(&pt->bo->ttm.base.gpuva.list));
>  	xe_bo_unpin(pt->bo);
> -	xe_bo_put_deferred(pt->bo, deferred, NULL);
> +	xe_bo_put_deferred(pt->bo, deferred, &added);
> +	if (added) {
> +		xe_assert(pt->bo->vm->xe, !kref_read(&pt->bo->ttm.base.refcount));
> +
> +		/*
> +		 * We need the VM present until the BO is destroyed as it shares
> +		 * a dma-resv and BO destroy is async. Reinit BO refcount so
> +		 * xe_bo_put_async can be used when the PT job ops refcount goes
> +		 * to zero.
> +		 */
> +		xe_vm_get(pt->bo->vm);
> +		pt->bo->flags |= XE_BO_FLAG_PUT_VM_ASYNC;
> +	}

[Severity: High]
Does this code forget to actually re-initialize the BO refcount? The comment
explicitly mentions reinitializing the BO refcount so xe_bo_put_async can be
used, but it looks like only the VM is retrieved and the flag is set. Without a
call to reinitialize the kref (like kref_init for pt->bo->ttm.base.refcount), 
the refcount remains at zero.

[ ... ]

> +static void xe_pt_job_ops_destroy(struct kref *ref)
> +{
> +	struct xe_pt_job_ops *pt_job_ops =
> +		container_of(ref, struct xe_pt_job_ops, refcount);
> +	struct llist_node *freed;
> +	struct xe_bo *bo, *next;
> +
> +	xe_pt_update_ops_free(pt_job_ops->ops,
> +			      pt_job_ops->current_op);
> +
> +	freed = llist_del_all(&pt_job_ops->deferred);
> +	if (freed) {
> +		llist_for_each_entry_safe(bo, next, freed, freed) {
> +			struct xe_bo_dev *bo_device =
> +				&xe_bo_device(bo)->bo_device;
> +			/*
> +			 * If called from run_job, we are in the dma-fencing
> +			 * path and cannot take dma-resv locks so use an async
> +			 * put.
> +			 */
> +			if (llist_add(&bo->freed, &bo_device->async_list))
> +				schedule_work(&bo_device->async_free);
> +		}
> +	}

[Severity: High]
Because the BO refcount was left at zero during xe_pt_destroy, will the later
xe_bo_put in the async_free worker trigger a refcount underflow here? A kref_put
on a zero refcount returns false and skips the release function (such as
drm_gem_object_free). Can this cause the page table BO memory to be permanently
leaked?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260904211613.3934307-1-matthew.brost@intel.com?part=9

  reply	other threads:[~2026-09-04 21:40 UTC|newest]

Thread overview: 48+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 21:15 [PATCH v6 00/24] CPU binds and ULLS on migration queue Matthew Brost
2026-09-04 21:15 ` [PATCH v6 01/24] drm/xe: Drop struct xe_migrate_pt_update argument from populate/clear vfuns Matthew Brost
2026-09-04 21:15 ` [PATCH v6 02/24] drm/xe: Add xe_migrate_update_pgtables_cpu_execute helper Matthew Brost
2026-09-04 21:28   ` sashiko-bot
2026-09-04 21:15 ` [PATCH v6 03/24] drm/xe: Decouple exec queue idle check from LRC Matthew Brost
2026-09-04 21:15 ` [PATCH v6 04/24] drm/xe: Add job count to GuC exec queue snapshot Matthew Brost
2026-09-04 21:23   ` sashiko-bot
2026-09-04 21:15 ` [PATCH v6 05/24] drm/xe: Update xe_bo_put_deferred arguments to include writeback flag Matthew Brost
2026-09-04 21:15 ` [PATCH v6 06/24] drm/xe: Add XE_BO_FLAG_PUT_VM_ASYNC Matthew Brost
2026-09-04 21:33   ` sashiko-bot
2026-09-11 13:10   ` Francois Dugast
2026-09-11 19:54     ` Matthew Brost
2026-09-12  0:27       ` Matthew Brost
2026-09-04 21:15 ` [PATCH v6 07/24] drm/xe: Update scheduler job layer to support PT jobs Matthew Brost
2026-09-04 21:37   ` sashiko-bot
2026-09-11 15:24   ` Francois Dugast
2026-09-11 19:25     ` Matthew Brost
2026-09-04 21:15 ` [PATCH v6 08/24] drm/xe: Add helpers to access PT ops Matthew Brost
2026-09-04 21:15 ` [PATCH v6 09/24] drm/xe: Add struct xe_pt_job_ops Matthew Brost
2026-09-04 21:40   ` sashiko-bot [this message]
2026-09-04 21:15 ` [PATCH v6 10/24] drm/xe: Update GuC submission backend to run PT jobs Matthew Brost
2026-09-04 21:39   ` sashiko-bot
2026-09-04 21:16 ` [PATCH v6 11/24] drm/xe: Store level in struct xe_vm_pgtable_update Matthew Brost
2026-09-04 21:16 ` [PATCH v6 12/24] drm/xe: Don't use migrate exec queue for page fault binds Matthew Brost
2026-09-04 21:16 ` [PATCH v6 13/24] drm/xe: Enable CPU binds for jobs Matthew Brost
2026-09-04 21:44   ` sashiko-bot
2026-09-04 21:16 ` [PATCH v6 14/24] drm/xe: Remove unused arguments from xe_migrate_pt_update_ops Matthew Brost
2026-09-04 21:16 ` [PATCH v6 15/24] drm/xe: Make bind queues operate cross-tile Matthew Brost
2026-09-04 21:16 ` [PATCH v6 16/24] drm/xe: Add CPU bind layer Matthew Brost
2026-09-04 21:50   ` sashiko-bot
2026-09-04 21:16 ` [PATCH v6 17/24] drm/xe: Add device flag to enable PT mirroring across tiles Matthew Brost
2026-09-04 21:40   ` sashiko-bot
2026-09-04 21:16 ` [PATCH v6 18/24] drm/xe: Add ULLS support to LRC Matthew Brost
2026-09-04 21:16 ` [PATCH v6 19/24] drm/xe: Add ULLS migration job support to migration layer Matthew Brost
2026-09-04 21:40   ` sashiko-bot
2026-09-04 21:16 ` [PATCH v6 20/24] drm/xe: Add ULLS migration job support to ring ops Matthew Brost
2026-09-04 21:16 ` [PATCH v6 21/24] drm/xe: Add ULLS migration job support to GuC submission Matthew Brost
2026-09-04 21:16 ` [PATCH v6 22/24] drm/xe: Enter ULLS for migration jobs upon page fault or SVM prefetch Matthew Brost
2026-09-04 21:16 ` [PATCH v6 23/24] drm/xe: Add modparam to enable / disable ULLS on migrate queue Matthew Brost
2026-09-09  8:03   ` Thomas Hellström
2026-09-09 18:11     ` Matthew Brost
2026-09-04 21:16 ` [PATCH v6 24/24] drm/xe: Document ULLS for migration jobs Matthew Brost
2026-09-09  9:01   ` Thomas Hellström
2026-09-09 17:53     ` Matthew Brost
2026-09-04 21:24 ` ✗ CI.checkpatch: warning for CPU binds and ULLS on migration queue (rev8) Patchwork
2026-09-04 21:26 ` ✓ CI.KUnit: success " Patchwork
2026-09-04 22:16 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-05  3:34 ` ✗ Xe.CI.FULL: failure " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260904214003.09F171F00A3D@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox