Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>
To: intel-xe@lists.freedesktop.org
Cc: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
	"Matthew Brost" <matthew.brost@intel.com>,
	"Rodrigo Vivi" <rodrigo.vivi@intel.com>,
	"Matthew Auld" <matthew.auld@intel.com>
Subject: [PATCH v4 0/3] drm/xe: Protect against premature module unloads
Date: Tue, 29 Sep 2026 16:29:07 +0200	[thread overview]
Message-ID: <20260929142910.47480-1-thomas.hellstrom@linux.intel.com> (raw)

Driver code is increasingly relying on bare drm_device references
(drm_dev_get()/drm_dev_put()) to keep a device's software state
around, without also pairing that with a reference on the owning
kernel module. Xe itself does this in several places, and so does
drm_gpuvm (used by xe) for the lifetime of a GPU VM. None of these
references currently prevent the xe module from being unloaded while
they, or the teardown work they can still trigger, are outstanding,
meaning driver code can end up executing after the module's own text
has already been freed.

This series closes that gap for xe:

- Patch 1 makes xe hold up module unload until every xe_device
  instance has actually been released, rather than only until the
  module's own refcount happens to reach zero, with a diagnostic if
  this ends up taking an unexpectedly long time.

- Patch 2 fixes a related, previously unprotected case where the
  teardown of a GPU VM or its address space mappings can be deferred
  to run at an arbitrary later time, including after module unload has
  already completed.

- Patch 3 fixes the same class of problem for execlist exec queue
  teardown, which could likewise be deferred past module unload.

Together, these changes ensure `rmmod xe` cannot free the module's
memory while any of its devices, or asynchronous work stemming from
them, might still be executing.

v2:
- Updated the commit message of patch 1 (then patch 2) to describe
  the actual implementation (a single 20s wait_var_event_timeout()
  followed by one pr_warn() and an unbounded wait_var_event(), rather
  than a loop retrying with a diagnostic every 10s) and its
  uninterruptible-sleep tradeoff (sashiko)

v3:
- drm_dev_release_barrier() now uses a single, global SRCU domain
  shared by all drivers instead of requiring each driver to register
  its own via a new &drm_driver.release_srcu field, accepting that a
  driver's call may then occasionally block on unrelated drivers'
  release callbacks (patch 1, patch 2, Christian König). This entire
  patch was later dropped in v4.

v4:
- Added a patch fixing a similar unprotected teardown gap in execlist
  exec queue destruction
- Dropped the core DRM patch that provided a global SRCU-based
  drm_dev_release_barrier() and the corresponding call to it in xe's
  module-exit path, in favor of relying solely on the xe_device
  instance count already tracked by this series
- Add code comments (patch 2)
- Use the xe->destroy_wq for vm->destroy_work (patch 2, Matt Brost)

Thomas Hellström (3):
  drm/xe: Don't unload the driver until all drm devices are freed
  drm/xe: Route deferred xe_vma/xe_vm teardown off system_dfl_wq
  drm/xe: Route execlist exec queue teardown off system_dfl_wq

 drivers/gpu/drm/xe/xe_device.c   | 29 +++++++++++++++++++++++++++++
 drivers/gpu/drm/xe/xe_device.h   |  2 ++
 drivers/gpu/drm/xe/xe_execlist.c |  6 +++++-
 drivers/gpu/drm/xe/xe_module.c   | 24 +++++++++++++++++++++---
 drivers/gpu/drm/xe/xe_vm.c       | 18 +++++++++++++++---
 5 files changed, 72 insertions(+), 7 deletions(-)

-- 
2.55.0


             reply	other threads:[~2026-09-29 14:29 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 14:29 Thomas Hellström [this message]
2026-09-29 14:29 ` [PATCH v4 1/3] drm/xe: Don't unload the driver until all drm devices are freed Thomas Hellström
2026-09-29 14:41   ` sashiko-bot
2026-09-29 14:51     ` Thomas Hellström
2026-09-29 14:29 ` [PATCH v4 2/3] drm/xe: Route deferred xe_vma/xe_vm teardown off system_dfl_wq Thomas Hellström
2026-09-29 14:29 ` [PATCH v4 3/3] drm/xe: Route execlist exec queue " Thomas Hellström
2026-09-29 16:18   ` Matthew Brost
2026-09-29 14:38 ` ✓ CI.KUnit: success for drm/xe: Protect against premature module unloads Patchwork
2026-09-29 16:04 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-29 18:14 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-10-01 10:06   ` Thomas Hellström

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929142910.47480-1-thomas.hellstrom@linux.intel.com \
    --to=thomas.hellstrom@linux.intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.auld@intel.com \
    --cc=matthew.brost@intel.com \
    --cc=rodrigo.vivi@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox