From: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>
To: intel-xe@lists.freedesktop.org
Cc: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
"Matthew Brost" <matthew.brost@intel.com>,
"Rodrigo Vivi" <rodrigo.vivi@intel.com>,
"Matthew Auld" <matthew.auld@intel.com>,
dri-devel@lists.freedesktop.org,
"Danilo Krummrich" <dakr@kernel.org>,
"Alice Ryhl" <aliceryhl@google.com>,
"Alex Deucher" <alexander.deucher@amd.com>,
"Christian König" <christian.koenig@amd.com>
Subject: [PATCH v2 0/3] drm, drm/xe: Protect against premature module unloads
Date: Thu, 24 Sep 2026 10:04:52 +0200 [thread overview]
Message-ID: <20260924080455.25458-1-thomas.hellstrom@linux.intel.com> (raw)
Driver and shared DRM helper code is increasingly relying on bare
drm_device references (drm_dev_get()/drm_dev_put()) to keep a device's
software state around, without also pairing that with a reference on
the owning kernel module. Xe itself does this in several places, and
so does drm_gpuvm for the lifetime of a GPU VM. None of these
references currently prevent the owning module from being unloaded
while they, or the teardown work they can still trigger, are
outstanding, meaning driver code can end up executing after its own
module's text has already been freed.
This series closes that gap for xe:
- Patch 1 adds core DRM infrastructure allowing a driver to wait for
its outstanding device-release callbacks to finish before
proceeding with module unload.
- Patch 2 makes xe use this infrastructure to hold up module unload
until every xe_device instance has actually been released, rather
than only until the module's own refcount happens to reach zero,
with a diagnostic if this ends up taking an unexpectedly long time.
- Patch 3 fixes a related, previously unprotected case where the
teardown of a GPU VM or its address space mappings can be deferred
to run at an arbitrary later time, including after module unload has
already completed.
Together, these changes ensure `rmmod xe` cannot free the module's
memory while any of its devices, or asynchronous work stemming from
them, might still be executing.
v2:
- Use plain WARN_ON_ONCE() instead of drm_WARN_ON_ONCE(NULL, ...) in
drm_dev_release_barrier(), fixing a NULL pointer dereference on the
warning path itself (patch 1, sashiko)
- Updated the commit message of patch 2 to describe the actual
implementation (a single 20s wait_var_event_timeout() followed by
one pr_warn() and an unbounded wait_var_event(), rather than a loop
retrying with a diagnostic every 10s) and its uninterruptible-sleep
tradeoff (sashiko)
Thomas Hellström (3):
drm: Provide a drm_dev_release_barrier() function to wait for device
release callbacks
drm/xe: Don't unload the driver until all drm devices are freed
drm/xe: Route deferred xe_vma/xe_vm teardown off system_dfl_wq
drivers/gpu/drm/drm_drv.c | 56 ++++++++++++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_device.c | 42 +++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_device.h | 2 ++
drivers/gpu/drm/xe/xe_module.c | 24 +++++++++++++--
drivers/gpu/drm/xe/xe_vm.c | 5 +--
include/drm/drm_drv.h | 24 +++++++++++++++
6 files changed, 148 insertions(+), 5 deletions(-)
--
2.55.0
next reply other threads:[~2026-09-24 8:05 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 8:04 Thomas Hellström [this message]
2026-09-24 8:04 ` [PATCH v2 1/3] drm: Provide a drm_dev_release_barrier() function to wait for device release callbacks Thomas Hellström
2026-09-24 11:19 ` Christian König
2026-09-24 13:10 ` Thomas Hellström
2026-09-24 8:04 ` [PATCH v2 2/3] drm/xe: Don't unload the driver until all drm devices are freed Thomas Hellström
2026-09-24 21:00 ` Matthew Brost
2026-09-24 8:04 ` [PATCH v2 3/3] drm/xe: Route deferred xe_vma/xe_vm teardown off system_dfl_wq Thomas Hellström
2026-09-24 20:55 ` Matthew Brost
2026-09-24 8:14 ` ✓ CI.KUnit: success for drm, drm/xe: Protect against premature module unloads (rev2) Patchwork
2026-09-24 9:04 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-24 22:28 ` ✓ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260924080455.25458-1-thomas.hellstrom@linux.intel.com \
--to=thomas.hellstrom@linux.intel.com \
--cc=alexander.deucher@amd.com \
--cc=aliceryhl@google.com \
--cc=christian.koenig@amd.com \
--cc=dakr@kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=matthew.auld@intel.com \
--cc=matthew.brost@intel.com \
--cc=rodrigo.vivi@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox