From: Zhu Lingshan <lingshan.zhu@amd.com>
To: <Alexander.Deucher@amd.com>, <Christian.Koenig@amd.com>,
<felix.kuehling@amd.com>
Cc: <Ray.Huang@amd.com>, <amd-gfx@lists.freedesktop.org>,
Zhu Lingshan <lingshan.zhu@amd.com>
Subject: [PATCH 07/10] drm/amdgpu: hold userq kref during isolation scheduling
Date: Fri, 28 Aug 2026 17:53:46 +0800 [thread overview]
Message-ID: <20260828095349.9797-8-lingshan.zhu@amd.com> (raw)
In-Reply-To: <20260828095349.9797-1-lingshan.zhu@amd.com>
The enforce-isolation stop and start handlers
iterate the doorbell XArray, but process a user queue
without holding its kref. Therefore, a concurrent queue
destruction process can free a queue before these paths
finish their work, causing use-after-free problems.
This commit fixes this issue by introducing a new helper
amdgpu_userq_xa_find() which finds a queue from a XArray
and hold its kref, and employ this helper in the
enforce-isolation stop and start handlers.
Signed-off-by: Zhu Lingshan <lingshan.zhu@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 72 +++++++++++++++++++++--
1 file changed, 68 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index 9fe20cb9af58..1427ff175dab 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -661,6 +661,54 @@ amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell)
return queue;
}
+/**
+ * amdgpu_userq_xa_find - search the XArray for a queue
+ * @xa: user queue XArray
+ * @index: first index to search, updated by every iteration of the search
+ *
+ * Return: a queue which has the lowest index that is at least @index,
+ * or NULL when no queue was found.
+ *
+ * The caller must release the kref of a queue with amdgpu_userq_put() after use.
+ */
+static struct amdgpu_usermode_queue *
+amdgpu_userq_xa_find(struct xarray *xa, unsigned long *index)
+{
+ struct amdgpu_usermode_queue *queue;
+ unsigned long flags;
+
+ xa_lock_irqsave(xa, flags);
+ queue = xa_find(xa, index, ULONG_MAX, XA_PRESENT);
+ while (queue) {
+ /*
+ * If found a queue but failed to get a kref,
+ * it means the queue is in destruction process,
+ * so skip it by continuing the loop.
+ *
+ * If get a kref of the queue, break the loop and return it.
+ */
+ if (kref_get_unless_zero(&queue->refcount))
+ break;
+
+ /*
+ * If the index is ULONG_MAX, we have reached the end of the XArray.
+ * Break the loop and return NULL because ULONG_MAX + 1 = 0,
+ * which is the start of the XArray.
+ */
+ if (*index == ULONG_MAX) {
+ queue = NULL;
+ break;
+ }
+
+ (*index)++;
+ queue = xa_find(xa, index, ULONG_MAX, XA_PRESENT);
+ }
+
+ xa_unlock_irqrestore(xa, flags);
+
+ return queue;
+}
+
void amdgpu_userq_put(struct amdgpu_usermode_queue *queue)
{
if (queue)
@@ -1496,7 +1544,7 @@ int amdgpu_userq_stop_sched_for_enforce_isolation(struct amdgpu_device *adev,
u32 ip_mask = amdgpu_userq_get_supported_ip_mask(adev);
struct amdgpu_usermode_queue *queue;
struct amdgpu_userq_mgr *uqm;
- unsigned long queue_id;
+ unsigned long queue_id = 0;
int ret = 0, r;
/* only need to stop gfx/compute */
@@ -1506,7 +1554,8 @@ int amdgpu_userq_stop_sched_for_enforce_isolation(struct amdgpu_device *adev,
if (adev->userq_halt_for_enforce_isolation)
dev_warn(adev->dev, "userq scheduling already stopped!\n");
adev->userq_halt_for_enforce_isolation = true;
- xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
+ queue = amdgpu_userq_xa_find(&adev->userq_doorbell_xa, &queue_id);
+ while (queue) {
uqm = queue->userq_mgr;
cancel_delayed_work_sync(&uqm->resume_work);
mutex_lock(&uqm->userq_mutex);
@@ -1518,6 +1567,13 @@ int amdgpu_userq_stop_sched_for_enforce_isolation(struct amdgpu_device *adev,
ret = r;
}
mutex_unlock(&uqm->userq_mutex);
+ amdgpu_userq_put(queue);
+
+ if (queue_id == ULONG_MAX)
+ break;
+ queue_id++;
+ queue = amdgpu_userq_xa_find(&adev->userq_doorbell_xa,
+ &queue_id);
}
return ret;
@@ -1529,7 +1585,7 @@ int amdgpu_userq_start_sched_for_enforce_isolation(struct amdgpu_device *adev,
u32 ip_mask = amdgpu_userq_get_supported_ip_mask(adev);
struct amdgpu_usermode_queue *queue;
struct amdgpu_userq_mgr *uqm;
- unsigned long queue_id;
+ unsigned long queue_id = 0;
int ret = 0, r;
/* only need to stop gfx/compute */
@@ -1541,7 +1597,8 @@ int amdgpu_userq_start_sched_for_enforce_isolation(struct amdgpu_device *adev,
adev->userq_halt_for_enforce_isolation = false;
- xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
+ queue = amdgpu_userq_xa_find(&adev->userq_doorbell_xa, &queue_id);
+ while (queue) {
uqm = queue->userq_mgr;
mutex_lock(&uqm->userq_mutex);
if (((queue->queue_type == AMDGPU_HW_IP_GFX) ||
@@ -1552,6 +1609,13 @@ int amdgpu_userq_start_sched_for_enforce_isolation(struct amdgpu_device *adev,
ret = r;
}
mutex_unlock(&uqm->userq_mutex);
+ amdgpu_userq_put(queue);
+
+ if (queue_id == ULONG_MAX)
+ break;
+ queue_id++;
+ queue = amdgpu_userq_xa_find(&adev->userq_doorbell_xa,
+ &queue_id);
}
return ret;
--
2.53.0
next prev parent reply other threads:[~2026-08-28 9:54 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 9:53 [PATCH 00/10] drm/amdgpu: secure userq lifecycle by its kref Zhu Lingshan
2026-08-28 9:53 ` [PATCH 01/10] drm/amdgpu: introduce amdgpu_lookup_queue_by_doorbell Zhu Lingshan
2026-08-28 13:08 ` Christian König
2026-08-28 15:59 ` Zhu, Lingshan
2026-08-28 9:53 ` [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues Zhu Lingshan
2026-08-28 13:09 ` Christian König
2026-08-28 15:59 ` Zhu, Lingshan
2026-08-28 16:26 ` Christian König
2026-08-28 9:53 ` [PATCH 03/10] drm/amdgpu/gfx11: hold userq refs in private fault worker Zhu Lingshan
2026-08-28 13:11 ` Christian König
2026-08-28 15:59 ` Zhu, Lingshan
2026-08-28 9:53 ` [PATCH 04/10] drm/amdgpu/gfx12: " Zhu Lingshan
2026-08-28 9:53 ` [PATCH 05/10] drm/amdgpu: implement asynchronous userq destruction routine Zhu Lingshan
2026-08-28 9:53 ` [PATCH 06/10] drm/amdgpu: hold userq kref in MES reset Zhu Lingshan
2026-08-28 9:53 ` Zhu Lingshan [this message]
2026-08-28 9:53 ` [PATCH 08/10] drm/amdgpu: hold userq kref during suspend and resume Zhu Lingshan
2026-08-28 9:53 ` [PATCH 09/10] drm/amdgpu: free userq by kref_put when fails to create Zhu Lingshan
2026-08-28 9:53 ` [PATCH 10/10] drm/amdgpu: take queue kref in userq_create to avoid UAF Zhu Lingshan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260828095349.9797-8-lingshan.zhu@amd.com \
--to=lingshan.zhu@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Ray.Huang@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=felix.kuehling@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.