Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Brost <matthew.brost@intel.com>
To: Arvind Yadav <arvind.yadav@intel.com>
Cc: <intel-xe@lists.freedesktop.org>,
	<dri-devel@lists.freedesktop.org>, <rodrigo.vivi@intel.com>,
	<himal.prasad.ghimiray@intel.com>,
	<thomas.hellstrom@linux.intel.com>
Subject: Re: [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
Date: Thu, 17 Sep 2026 20:22:33 -0700	[thread overview]
Message-ID: <aqyueX3xWoVJ6P8b@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260916095337.3104891-2-arvind.yadav@intel.com>

On Wed, Sep 16, 2026 at 03:23:33PM +0530, Arvind Yadav wrote:
> xe_vm_free() is the drm_gpuvm vm_free callback. It hands the final
> teardown to vm_destroy_work_func() on a workqueue and returns.
> 
> drm_gpuvm_free() drops its device reference immediately after the
> callback returns:
> 
> 	gpuvm->ops->vm_free(gpuvm);
> 	drm_dev_put(drm);
> 
> vm_destroy_work_func() then keeps using device state: xe_pm_runtime_put()
> for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the tile
> iteration. If the freed VM held the last device reference, the work runs
> against a released xe_device.
> 
> Take a device reference in xe_vm_free() and drop it once
> vm_destroy_work_func() has finished using the device.
> 
> Cc: Matthew Brost <matthew.brost@intel.com>

This is a fix, IMO. Ideally, we should probably push the delayed-destroy
semantics into gpuvm if they are really needed. I'm also questioning
whether the VM destroy worker is actually required. This dates back to
the very early days of Xe, and I doubt we've ever revisited whether it
is necessary.

Let's follow up with one of the following:
- Introduce async destroy in gpuvm and have it own the drm_dev_get/put.
- Drop delayed destroy entirely in Xe.

As a temporary fix that can be backported, this looks good to me, so
with a Fixes tag:

Reviewed-by: Matthew Brost <matthew.brost@intel.com>

> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> Assisted-by: Claude:claude-opus-4-8
> Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> ---
>  drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
>  1 file changed, 9 insertions(+)
> 
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index efa5ff6cc823..264bdab75de2 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
> @@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct work_struct *w)
>  		xe_file_put(vm->xef);
>  
>  	kfree(vm);
> +
> +	drm_dev_put(&xe->drm);
>  }
>  
>  static void xe_vm_free(struct drm_gpuvm *gpuvm)
>  {
>  	struct xe_vm *vm = container_of(gpuvm, struct xe_vm, gpuvm);
>  
> +	/*
> +	 * drm_gpuvm drops its device reference as soon as this callback
> +	 * returns, but vm_destroy_work_func() still uses device state. Hold a
> +	 * reference across the deferred work.
> +	 */
> +	drm_dev_get(&vm->xe->drm);
> +
>  	/* To destroy the VM we need to be able to sleep */
>  	queue_work(system_dfl_wq, &vm->destroy_work);
>  }
> -- 
> 2.43.0
> 

  reply	other threads:[~2026-09-18  3:22 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
2026-09-16  9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
2026-09-18  3:22   ` Matthew Brost [this message]
2026-09-18  7:22     ` Thomas Hellström
2026-09-18 22:41       ` Matthew Brost
2026-09-21  7:06         ` Thomas Hellström
2026-09-16  9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
2026-09-18 22:36   ` Matthew Brost
2026-09-21  6:46     ` Yadav, Arvind
2026-09-16  9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
2026-09-16  9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
2026-09-16  9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
2026-09-16 10:01 ` ✓ CI.KUnit: success for drm/xe: Fix VM teardown and migration queue recovery Patchwork
2026-09-16 10:59 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-16 12:11 ` ✓ Xe.CI.FULL: " Patchwork
2026-09-18 22:31 ` [PATCH 0/5] " Matthew Brost
2026-09-24 10:06   ` Yadav, Arvind

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqyueX3xWoVJ6P8b@gsse-cloud1.jf.intel.com \
    --to=matthew.brost@intel.com \
    --cc=arvind.yadav@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=himal.prasad.ghimiray@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=rodrigo.vivi@intel.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox