From: "Jonghyuk Kim(MalHyuk)" <malhyuk97@gmail.com>
To: tursulin@ursulin.net, phasta@kernel.org, matthew.brost@intel.com,
dakr@kernel.org
Cc: christian.koenig@amd.com, dri-devel@lists.freedesktop.org,
linux-kernel@vger.kernel.org,
"Jonghyuk Kim(MalHyuk)" <malhyuk97@gmail.com>
Subject: [PATCH v3 1/2] drm/sched: fix use-after-free of the fence timeline name
Date: Wed, 2 Sep 2026 23:42:03 +0900 [thread overview]
Message-ID: <20260902144204.1843670-2-malhyuk97@gmail.com> (raw)
In-Reply-To: <20260902144204.1843670-1-malhyuk97@gmail.com>
drm_sched_fence_get_timeline_name() returns fence->sched->name. A driver
that allocates a drm_gpu_scheduler per context, queue or VM frees that
scheduler when the context is destroyed, but the finished fence can
outlive it: unprivileged userspace can hold the exported fence through a
sync_file or drm_syncobj and query its timeline name afterwards (e.g. via
SYNC_IOC_FILE_INFO), dereferencing the freed scheduler. This is a
slab-use-after-free read; once the slab is re-sprayed it becomes an info
leak. It is the same class as CVE-2025-38703 (drm/xe) and CVE-2025-71302
(drm/panthor), which were fixed per driver. amdxdna, nouveau and msm
(VM_BIND) allocate per-context schedulers and are still affected.
The dma-fence contract already forbids touching driver-provided data -
the memory reachable through &dma_fence.ops - once a fence is signalled,
and dma_fence_timeline_name() enforces it: after the ops are detached it
returns a static string instead of calling get_timeline_name().
dma_fence detaches the ops on signalling, but only for fences that carry
neither a .release nor a .wait callback (see
dma_fence_signal_timestamp_locked()).
The finished fence carried a .release callback solely to drop the
scheduled fence's reference. That callback kept the ops attached, leaving
get_timeline_name() reachable on a signalled finished fence with a
dangling ->sched. Drop the callback and move the reference handling
instead:
- The scheduled fence now holds a reference on the finished fence, so
the finished fence, and with it the shared allocation, is released
last. Its release drops the parent fence and that finished-fence
reference; the finished fence is then freed from dma_fence_free().
This requires @finished to sit at offset 0 of struct drm_sched_fence,
since dma_fence_free() ultimately kfree()s the fence pointer.
- drm_sched_job_cleanup() drops the scheduled fence's initial reference,
which the removed .release used to cascade.
With the finished fence no longer carrying .release its ops are detached
on signalling, so get_timeline_name() can no longer run against a freed
scheduler. Unlike caching the name string, this also covers drivers whose
timeline name is dynamically allocated (drm/panthor, drm/xe).
Detaching the ops on signalling also makes to_drm_sched_fence() return
NULL for a signalled finished fence. Callers already handle a NULL return
- it is the normal result for a foreign fence - and a signalled fence is
an already-satisfied dependency, so the scheduler's dependency-collapsing
optimisation is unaffected. It additionally avoids the container_of() on
a possibly-freed foreign scheduler that amdgpu_sync_same_dev() and
pvr_queue_fence_is_native() would otherwise perform.
Suggested-by: Philipp Stanner <phasta@kernel.org>
Signed-off-by: Jonghyuk Kim(MalHyuk) <malhyuk97@gmail.com>
---
drivers/gpu/drm/scheduler/sched_fence.c | 46 ++++++++--------
drivers/gpu/drm/scheduler/sched_main.c | 9 ++++
include/drm/gpu_scheduler.h | 72 ++++++++++++++-----------
3 files changed, 71 insertions(+), 56 deletions(-)
diff --git a/drivers/gpu/drm/scheduler/sched_fence.c b/drivers/gpu/drm/scheduler/sched_fence.c
index 096fe28aa9c9..f463afa0ee4e 100644
--- a/drivers/gpu/drm/scheduler/sched_fence.c
+++ b/drivers/gpu/drm/scheduler/sched_fence.c
@@ -95,15 +95,6 @@ static const char *drm_sched_fence_get_timeline_name(struct dma_fence *f)
return (const char *)fence->sched->name;
}
-static void drm_sched_fence_free_rcu(struct rcu_head *rcu)
-{
- struct dma_fence *f = container_of(rcu, struct dma_fence, rcu);
- struct drm_sched_fence *fence = to_drm_sched_fence(f);
-
- if (!WARN_ON_ONCE(!fence))
- kmem_cache_free(sched_fence_slab, fence);
-}
-
/**
* drm_sched_fence_free - free up an uninitialized fence
*
@@ -132,21 +123,12 @@ static void drm_sched_fence_release_scheduled(struct dma_fence *f)
struct drm_sched_fence *fence = to_drm_sched_fence(f);
dma_fence_put(fence->parent);
- call_rcu(&fence->finished.rcu, drm_sched_fence_free_rcu);
-}
-
-/**
- * drm_sched_fence_release_finished - drop extra reference
- *
- * @f: fence
- *
- * Drop the extra reference from the scheduled fence to the base fence.
- */
-static void drm_sched_fence_release_finished(struct dma_fence *f)
-{
- struct drm_sched_fence *fence = to_drm_sched_fence(f);
-
- dma_fence_put(&fence->scheduled);
+ /*
+ * Drop the reference the scheduled fence holds on the finished fence.
+ * The finished fence is released last and frees the shared allocation
+ * from its dma_fence_free() (see drm_sched_fence_init()).
+ */
+ dma_fence_put(&fence->finished);
}
static void drm_sched_fence_set_deadline_finished(struct dma_fence *f,
@@ -189,7 +171,13 @@ static const struct dma_fence_ops drm_sched_fence_ops_scheduled = {
static const struct dma_fence_ops drm_sched_fence_ops_finished = {
.get_driver_name = drm_sched_fence_get_driver_name,
.get_timeline_name = drm_sched_fence_get_timeline_name,
- .release = drm_sched_fence_release_finished,
+ /*
+ * No .release callback: dma_fence detaches ->ops on signalling for
+ * fences without .release/.wait, so get_timeline_name() is never called
+ * on a signalled finished fence and cannot dereference a freed
+ * scheduler. The shared allocation is freed from dma_fence_free() once
+ * this fence's refcount drops - it is released last, after @scheduled.
+ */
.set_deadline = drm_sched_fence_set_deadline_finished,
};
@@ -233,6 +221,14 @@ void drm_sched_fence_init(struct drm_sched_fence *fence,
&fence->lock, entity->fence_context, seq);
dma_fence_init(&fence->finished, &drm_sched_fence_ops_finished,
&fence->lock, entity->fence_context + 1, seq);
+
+ /*
+ * Hold a reference on the finished fence from the scheduled fence, so
+ * the finished fence (and the shared allocation) outlives @scheduled.
+ * drm_sched_fence_release_scheduled() drops it; the finished fence is
+ * therefore released last and frees the allocation via dma_fence_free().
+ */
+ dma_fence_get(&fence->finished);
}
module_init(drm_sched_fence_slab_init);
diff --git a/drivers/gpu/drm/scheduler/sched_main.c b/drivers/gpu/drm/scheduler/sched_main.c
index 6cb6f9546493..fb238f51c0ed 100644
--- a/drivers/gpu/drm/scheduler/sched_main.c
+++ b/drivers/gpu/drm/scheduler/sched_main.c
@@ -842,6 +842,15 @@ void drm_sched_job_cleanup(struct drm_sched_job *job)
* been called.
*/
dma_fence_put(&job->s_fence->finished);
+ /*
+ * Drop the initial reference on the scheduled fence. It no
+ * longer has a .release callback dropping it (the finished
+ * fence's .release was removed to allow ops-detach on signal),
+ * so the last put here lets drm_sched_fence_release_scheduled()
+ * run, which drops @parent and the scheduled fence's reference
+ * on @finished. @finished is freed last, from dma_fence_free().
+ */
+ dma_fence_put(&job->s_fence->scheduled);
drm_sched_entity_stats_put(job->entity_stats);
} else {
/* The job was aborted before it has been committed to be run;
diff --git a/include/drm/gpu_scheduler.h b/include/drm/gpu_scheduler.h
index 7a64cc11de08..686c3687944f 100644
--- a/include/drm/gpu_scheduler.h
+++ b/include/drm/gpu_scheduler.h
@@ -287,48 +287,58 @@ struct drm_sched_rq {
* struct drm_sched_fence - fences corresponding to the scheduling of a job.
*/
struct drm_sched_fence {
- /**
- * @scheduled: this fence is what will be signaled by the scheduler
- * when the job is scheduled.
- */
- struct dma_fence scheduled;
-
- /**
- * @finished: this fence is what will be signaled by the scheduler
- * when the job is completed.
- *
- * When setting up an out fence for the job, you should use
- * this, since it's available immediately upon
- * drm_sched_job_init(), and the fence returned by the driver
- * from run_job() won't be created until the dependencies have
- * resolved.
- */
+ /**
+ * @finished: this fence is what will be signaled by the scheduler
+ * when the job is completed.
+ *
+ * When setting up an out fence for the job, you should use
+ * this, since it's available immediately upon
+ * drm_sched_job_init(), and the fence returned by the driver
+ * from run_job() won't be created until the dependencies have
+ * resolved.
+ *
+ * @finished is kept first in the struct: it is the fence exported to
+ * userspace and therefore the one whose &dma_fence_ops.release is
+ * dropped so that dma_fence detaches its ops on signalling. It is
+ * released last (see &drm_sched_fence.scheduled) and frees the whole
+ * object via dma_fence_free(), which requires it to sit at offset 0.
+ */
struct dma_fence finished;
+ /**
+ * @scheduled: this fence is what will be signaled by the scheduler
+ * when the job is scheduled.
+ *
+ * It holds a reference on @finished so that the shared allocation is
+ * released only after @scheduled itself is done; its release drops
+ * that reference and the @parent one.
+ */
+ struct dma_fence scheduled;
+
/**
* @deadline: deadline set on &drm_sched_fence.finished which
* potentially needs to be propagated to &drm_sched_fence.parent
*/
ktime_t deadline;
- /**
- * @parent: the fence returned by &drm_sched_backend_ops.run_job
- * when scheduling the job on hardware. We signal the
- * &drm_sched_fence.finished fence once parent is signalled.
- */
+ /**
+ * @parent: the fence returned by &drm_sched_backend_ops.run_job
+ * when scheduling the job on hardware. We signal the
+ * &drm_sched_fence.finished fence once parent is signalled.
+ */
struct dma_fence *parent;
- /**
- * @sched: the scheduler instance to which the job having this struct
- * belongs to.
- */
+ /**
+ * @sched: the scheduler instance to which the job having this struct
+ * belongs to.
+ */
struct drm_gpu_scheduler *sched;
- /**
- * @lock: the lock used by the scheduled and the finished fences.
- */
+ /**
+ * @lock: the lock used by the scheduled and the finished fences.
+ */
spinlock_t lock;
- /**
- * @owner: job owner for debugging
- */
+ /**
+ * @owner: job owner for debugging
+ */
void *owner;
/**
--
2.43.0
next prev parent reply other threads:[~2026-09-03 7:33 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 14:42 [PATCH v3 0/2] drm/sched: fix use-after-free of the fence timeline name Jonghyuk Kim(MalHyuk)
2026-09-02 14:42 ` Jonghyuk Kim(MalHyuk) [this message]
2026-09-02 14:58 ` [PATCH v3 1/2] " sashiko-bot
2026-09-03 8:46 ` Philipp Stanner
2026-09-03 10:02 ` Christian König
2026-09-03 10:22 ` Christian König
2026-09-02 14:42 ` [PATCH v3 2/2] drm/sched/tests: add a UAF regression test for the " Jonghyuk Kim(MalHyuk)
2026-09-02 15:04 ` sashiko-bot
2026-09-02 16:09 ` [PATCH v3 0/2] drm/sched: fix use-after-free of the fence " Philipp Stanner
2026-09-02 17:25 ` Philipp Stanner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260902144204.1843670-2-malhyuk97@gmail.com \
--to=malhyuk97@gmail.com \
--cc=christian.koenig@amd.com \
--cc=dakr@kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=linux-kernel@vger.kernel.org \
--cc=matthew.brost@intel.com \
--cc=phasta@kernel.org \
--cc=tursulin@ursulin.net \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox