From: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>
To: intel-xe@lists.freedesktop.org
Cc: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
"Matthew Brost" <matthew.brost@intel.com>,
"Rodrigo Vivi" <rodrigo.vivi@intel.com>,
"Matthew Auld" <matthew.auld@intel.com>
Subject: [PATCH v4 0/3] drm/xe: Protect against premature module unloads
Date: Tue, 29 Sep 2026 16:29:07 +0200 [thread overview]
Message-ID: <20260929142910.47480-1-thomas.hellstrom@linux.intel.com> (raw)
Driver code is increasingly relying on bare drm_device references
(drm_dev_get()/drm_dev_put()) to keep a device's software state
around, without also pairing that with a reference on the owning
kernel module. Xe itself does this in several places, and so does
drm_gpuvm (used by xe) for the lifetime of a GPU VM. None of these
references currently prevent the xe module from being unloaded while
they, or the teardown work they can still trigger, are outstanding,
meaning driver code can end up executing after the module's own text
has already been freed.
This series closes that gap for xe:
- Patch 1 makes xe hold up module unload until every xe_device
instance has actually been released, rather than only until the
module's own refcount happens to reach zero, with a diagnostic if
this ends up taking an unexpectedly long time.
- Patch 2 fixes a related, previously unprotected case where the
teardown of a GPU VM or its address space mappings can be deferred
to run at an arbitrary later time, including after module unload has
already completed.
- Patch 3 fixes the same class of problem for execlist exec queue
teardown, which could likewise be deferred past module unload.
Together, these changes ensure `rmmod xe` cannot free the module's
memory while any of its devices, or asynchronous work stemming from
them, might still be executing.
v2:
- Updated the commit message of patch 1 (then patch 2) to describe
the actual implementation (a single 20s wait_var_event_timeout()
followed by one pr_warn() and an unbounded wait_var_event(), rather
than a loop retrying with a diagnostic every 10s) and its
uninterruptible-sleep tradeoff (sashiko)
v3:
- drm_dev_release_barrier() now uses a single, global SRCU domain
shared by all drivers instead of requiring each driver to register
its own via a new &drm_driver.release_srcu field, accepting that a
driver's call may then occasionally block on unrelated drivers'
release callbacks (patch 1, patch 2, Christian König). This entire
patch was later dropped in v4.
v4:
- Added a patch fixing a similar unprotected teardown gap in execlist
exec queue destruction
- Dropped the core DRM patch that provided a global SRCU-based
drm_dev_release_barrier() and the corresponding call to it in xe's
module-exit path, in favor of relying solely on the xe_device
instance count already tracked by this series
- Add code comments (patch 2)
- Use the xe->destroy_wq for vm->destroy_work (patch 2, Matt Brost)
Thomas Hellström (3):
drm/xe: Don't unload the driver until all drm devices are freed
drm/xe: Route deferred xe_vma/xe_vm teardown off system_dfl_wq
drm/xe: Route execlist exec queue teardown off system_dfl_wq
drivers/gpu/drm/xe/xe_device.c | 29 +++++++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_device.h | 2 ++
drivers/gpu/drm/xe/xe_execlist.c | 6 +++++-
drivers/gpu/drm/xe/xe_module.c | 24 +++++++++++++++++++++---
drivers/gpu/drm/xe/xe_vm.c | 18 +++++++++++++++---
5 files changed, 72 insertions(+), 7 deletions(-)
--
2.55.0
next reply other threads:[~2026-09-29 14:29 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 14:29 Thomas Hellström [this message]
2026-09-29 14:29 ` [PATCH v4 1/3] drm/xe: Don't unload the driver until all drm devices are freed Thomas Hellström
2026-09-29 14:41 ` sashiko-bot
2026-09-29 14:51 ` Thomas Hellström
2026-09-29 14:29 ` [PATCH v4 2/3] drm/xe: Route deferred xe_vma/xe_vm teardown off system_dfl_wq Thomas Hellström
2026-09-29 14:29 ` [PATCH v4 3/3] drm/xe: Route execlist exec queue " Thomas Hellström
2026-09-29 16:18 ` Matthew Brost
2026-09-29 14:38 ` ✓ CI.KUnit: success for drm/xe: Protect against premature module unloads Patchwork
2026-09-29 16:04 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-29 18:14 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-10-01 10:06 ` Thomas Hellström
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929142910.47480-1-thomas.hellstrom@linux.intel.com \
--to=thomas.hellstrom@linux.intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=matthew.auld@intel.com \
--cc=matthew.brost@intel.com \
--cc=rodrigo.vivi@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox