Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>
To: Matthew Brost <matthew.brost@intel.com>,
	Arvind Yadav <arvind.yadav@intel.com>
Cc: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org,
	 rodrigo.vivi@intel.com, himal.prasad.ghimiray@intel.com
Subject: Re: [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
Date: Fri, 18 Sep 2026 09:22:53 +0200	[thread overview]
Message-ID: <2b196d015b37329cd6eda25c565c83c5e5f77a7f.camel@linux.intel.com> (raw)
In-Reply-To: <aqyueX3xWoVJ6P8b@gsse-cloud1.jf.intel.com>

On Thu, 2026-09-17 at 20:22 -0700, Matthew Brost wrote:
> On Wed, Sep 16, 2026 at 03:23:33PM +0530, Arvind Yadav wrote:
> > xe_vm_free() is the drm_gpuvm vm_free callback. It hands the final
> > teardown to vm_destroy_work_func() on a workqueue and returns.
> > 
> > drm_gpuvm_free() drops its device reference immediately after the
> > callback returns:
> > 
> > 	gpuvm->ops->vm_free(gpuvm);
> > 	drm_dev_put(drm);
> > 
> > vm_destroy_work_func() then keeps using device state:
> > xe_pm_runtime_put()
> > for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the
> > tile
> > iteration. If the freed VM held the last device reference, the work
> > runs
> > against a released xe_device.
> > 
> > Take a device reference in xe_vm_free() and drop it once
> > vm_destroy_work_func() has finished using the device.
> > 
> > Cc: Matthew Brost <matthew.brost@intel.com>
> 
> This is a fix, IMO. Ideally, we should probably push the delayed-
> destroy
> semantics into gpuvm if they are really needed. I'm also questioning
> whether the VM destroy worker is actually required. This dates back
> to
> the very early days of Xe, and I doubt we've ever revisited whether
> it
> is necessary.
> 
> Let's follow up with one of the following:
> - Introduce async destroy in gpuvm and have it own the
> drm_dev_get/put.
> - Drop delayed destroy entirely in Xe.

We need to keep in mind that the drm file keeps a reference on the Xe
module. So once the last close() callback has executed, the module can
typically be unloaded, causing execution UAF. It's therefore not really
recommended to keep file-related structures around with a refcount
after close.

Device references however typically don't necessarily keep the module
pinned. I had a series to fix this for xe only, (Got stalled) [1], but
in general we should be careful about leaking that assumption into DRM
code.

[1] https://patchwork.freedesktop.org/series/163298/

Thanks,
Thomas


> 
> As a temporary fix that can be backported, this looks good to me, so
> with a Fixes tag:
> 
> Reviewed-by: Matthew Brost <matthew.brost@intel.com>
> 
> > Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> > Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> > Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> > Assisted-by: Claude:claude-opus-4-8
> > Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> > ---
> >  drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
> >  1 file changed, 9 insertions(+)
> > 
> > diff --git a/drivers/gpu/drm/xe/xe_vm.c
> > b/drivers/gpu/drm/xe/xe_vm.c
> > index efa5ff6cc823..264bdab75de2 100644
> > --- a/drivers/gpu/drm/xe/xe_vm.c
> > +++ b/drivers/gpu/drm/xe/xe_vm.c
> > @@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct
> > work_struct *w)
> >  		xe_file_put(vm->xef);
> >  
> >  	kfree(vm);
> > +
> > +	drm_dev_put(&xe->drm);
> >  }
> >  
> >  static void xe_vm_free(struct drm_gpuvm *gpuvm)
> >  {
> >  	struct xe_vm *vm = container_of(gpuvm, struct xe_vm,
> > gpuvm);
> >  
> > +	/*
> > +	 * drm_gpuvm drops its device reference as soon as this
> > callback
> > +	 * returns, but vm_destroy_work_func() still uses device
> > state. Hold a
> > +	 * reference across the deferred work.
> > +	 */
> > +	drm_dev_get(&vm->xe->drm);
> > +
> >  	/* To destroy the VM we need to be able to sleep */
> >  	queue_work(system_dfl_wq, &vm->destroy_work);
> >  }
> > -- 
> > 2.43.0
> > 

  reply	other threads:[~2026-09-18  7:22 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
2026-09-16  9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
2026-09-18  3:22   ` Matthew Brost
2026-09-18  7:22     ` Thomas Hellström [this message]
2026-09-18 22:41       ` Matthew Brost
2026-09-21  7:06         ` Thomas Hellström
2026-09-16  9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
2026-09-18 22:36   ` Matthew Brost
2026-09-21  6:46     ` Yadav, Arvind
2026-09-16  9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
2026-09-16  9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
2026-09-16  9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
2026-09-16 10:01 ` ✓ CI.KUnit: success for drm/xe: Fix VM teardown and migration queue recovery Patchwork
2026-09-16 10:59 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-16 12:11 ` ✓ Xe.CI.FULL: " Patchwork
2026-09-18 22:31 ` [PATCH 0/5] " Matthew Brost
2026-09-24 10:06   ` Yadav, Arvind

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2b196d015b37329cd6eda25c565c83c5e5f77a7f.camel@linux.intel.com \
    --to=thomas.hellstrom@linux.intel.com \
    --cc=arvind.yadav@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=himal.prasad.ghimiray@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=rodrigo.vivi@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox