From: "Christian König" <christian.koenig@amd.com>
To: "Zhu, Lingshan" <lingshan.zhu@amd.com>,
Alexander.Deucher@amd.com, felix.kuehling@amd.com
Cc: Ray.Huang@amd.com, amd-gfx@lists.freedesktop.org
Subject: Re: [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues
Date: Fri, 28 Aug 2026 18:26:06 +0200 [thread overview]
Message-ID: <847308aa-e4cb-405f-9e9c-dda3e4d28eef@amd.com> (raw)
In-Reply-To: <7f53ac18-ecb5-4eaa-8692-67b9237447b2@amd.com>
On 8/28/26 17:59, Zhu, Lingshan wrote:
> On 8/28/2026 9:09 PM, Christian König wrote:
>
>> On 8/28/26 11:53, Zhu Lingshan wrote:
>>> The life cycle of a user queue is managed by its
>>> kref. However when destroy a userq manager,
>>> the kref_put of its queues in amdgpu_userq_mgr_fini
>>> may not be the last put, therefore the queues
>>> could be still alive after the userq manager
>>> has been destroyed, resulting in
>>> userq->userq_mgr use-after-free issues.
>>>
>>> This commit fixes this problem by introduce a new
>>> counter refs representing for the number of its queues,
>>> and only free the userq_manager when refs == 0
>> Clear NAK to that one as well, this is just nonsense.
>
> It could be better to have some explanations.
>
> I am not sure how to guarantee the put_kref in amdgpu_userq_mgr_fini
> is the last put and result in kref == 0, if not the last one,
> there can be userq->userq_mgr UAF bugs.
The rules are actually pretty simple:
The reference is for keeping the userq alive while IOCTLs happen. And IOCTL can only happen while the file and therefor the fpriv, userq_mgr etc... are still alive.
What can potentially be is that we also need to grab a reference from a work item, but in this case the fpriv/userq_mgr cleanup functions just need to cancel and wait for the work to finish.
There should *never* be a reference grabbed from interrupt context, explicitely because releasing that reference is also not possible from interrupt context. Instead xa_lock_irqsave() needs to be used to make sure that the userq stays alive while the interrupt processing happens.
Regards,
Christian.
>
> Thanks
> Lingshan
>
>> Christian.
>>
>>> Signed-off-by: Zhu Lingshan <lingshan.zhu@amd.com>
>>> ---
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 30 +++++++++++++++++++++++
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h | 9 +++++++
>>> 2 files changed, 39 insertions(+)
>>>
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>> index e0639f844a8e..f398986a61a5 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>> @@ -27,6 +27,7 @@
>>> #include <linux/pm_runtime.h>
>>> #include <linux/overflow.h>
>>> #include <drm/drm_drv.h>
>>> +#include <linux/wait_bit.h>
>>>
>>> #include "amdgpu.h"
>>> #include "amdgpu_reset.h"
>>> @@ -533,6 +534,17 @@ amdgpu_userq_get_doorbell_index(struct amdgpu_userq_mgr *uq_mgr,
>>> return r;
>>> }
>>>
>>> +static void amdgpu_userq_mgr_inc_refs(struct amdgpu_userq_mgr *uq_mgr)
>>> +{
>>> + atomic_inc(&uq_mgr->refs);
>>> +}
>>> +
>>> +static void amdgpu_userq_mgr_dec_refs(struct amdgpu_userq_mgr *uq_mgr)
>>> +{
>>> + if (atomic_dec_and_test(&uq_mgr->refs))
>>> + wake_up_var(&uq_mgr->refs);
>>> +}
>>> +
>>> static int
>>> amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_queue *queue)
>>> {
>>> @@ -594,6 +606,8 @@ static void amdgpu_userq_kref_destroy(struct kref *kref)
>>> r = amdgpu_userq_destroy(uq_mgr, queue);
>>> if (r)
>>> drm_file_err(uq_mgr->file, "Failed to destroy usermode queue %d\n", r);
>>> +
>>> + amdgpu_userq_mgr_dec_refs(uq_mgr);
>>> }
>>>
>>> struct amdgpu_usermode_queue *amdgpu_userq_get(struct amdgpu_userq_mgr *uq_mgr, u32 qid)
>>> @@ -707,6 +721,7 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
>>> queue->xcp_id = (fpriv->xcp_id != AMDGPU_XCP_NO_PARTITION) ?
>>> fpriv->xcp_id : 0;
>>> queue->userq_mgr = uq_mgr;
>>> + amdgpu_userq_mgr_inc_refs(uq_mgr);
>>> INIT_DELAYED_WORK(&queue->hang_detect_work,
>>> amdgpu_userq_hang_detect_work);
>>>
>>> @@ -819,6 +834,7 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
>>> free_queue:
>>> trace_amdgpu_userq_create_end(queue, r);
>>> kfree(queue);
>>> + amdgpu_userq_mgr_dec_refs(uq_mgr);
>>> err_pm_runtime:
>>> pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
>>> return r;
>>> @@ -1331,6 +1347,7 @@ int amdgpu_userq_mgr_init(struct amdgpu_userq_mgr *userq_mgr, struct drm_file *f
>>> {
>>> mutex_init(&userq_mgr->userq_mutex);
>>> xa_init_flags(&userq_mgr->userq_xa, XA_FLAGS_ALLOC);
>>> + atomic_set(&userq_mgr->refs, 0);
>>> userq_mgr->adev = adev;
>>> userq_mgr->file = file_priv;
>>> userq_mgr->proc_ctx_allocated = false;
>>> @@ -1380,6 +1397,19 @@ void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr)
>>> amdgpu_userq_put(queue);
>>> }
>>>
>>> + /*
>>> + * The above amdgpu_userq_put() may not be the last put
>>> + * of the kref of a user queue, therefore there could
>>> + * be some queues still alive even when the userq manager
>>> + * has been destroyed. This wait_evet() blocks
>>> + * amdgpu_userq_mgr_fini(), so keep userq_mgr alive
>>> + * while any queues holding it.
>>> + *
>>> + * This prevents queue->userq_mgr use-after-free issues.
>>> + */
>>> + wait_var_event(&userq_mgr->refs,
>>> + !atomic_read_acquire(&userq_mgr->refs));
>>> +
>>> xa_destroy(&userq_mgr->userq_xa);
>>>
>>> /*
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>> index 8fc73862f64e..a13d8d4dd5c7 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>> @@ -126,6 +126,15 @@ struct amdgpu_userq_mgr {
>>> */
>>> struct xarray userq_xa;
>>> struct mutex userq_mutex;
>>> +
>>> + /**
>>> + * @refs:
>>> + *
>>> + * Each queue increases this counter when join this manager,
>>> + * and decreases it when leave this manager.
>>> + */
>>> + atomic_t refs;
>>> +
>>> struct amdgpu_device *adev;
>>> struct delayed_work resume_work;
>>> struct drm_file *file;
next prev parent reply other threads:[~2026-08-28 16:26 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 9:53 [PATCH 00/10] drm/amdgpu: secure userq lifecycle by its kref Zhu Lingshan
2026-08-28 9:53 ` [PATCH 01/10] drm/amdgpu: introduce amdgpu_lookup_queue_by_doorbell Zhu Lingshan
2026-08-28 13:08 ` Christian König
2026-08-28 15:59 ` Zhu, Lingshan
2026-08-28 9:53 ` [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues Zhu Lingshan
2026-08-28 13:09 ` Christian König
2026-08-28 15:59 ` Zhu, Lingshan
2026-08-28 16:26 ` Christian König [this message]
2026-08-28 9:53 ` [PATCH 03/10] drm/amdgpu/gfx11: hold userq refs in private fault worker Zhu Lingshan
2026-08-28 13:11 ` Christian König
2026-08-28 15:59 ` Zhu, Lingshan
2026-08-28 9:53 ` [PATCH 04/10] drm/amdgpu/gfx12: " Zhu Lingshan
2026-08-28 9:53 ` [PATCH 05/10] drm/amdgpu: implement asynchronous userq destruction routine Zhu Lingshan
2026-08-28 9:53 ` [PATCH 06/10] drm/amdgpu: hold userq kref in MES reset Zhu Lingshan
2026-08-28 9:53 ` [PATCH 07/10] drm/amdgpu: hold userq kref during isolation scheduling Zhu Lingshan
2026-08-28 9:53 ` [PATCH 08/10] drm/amdgpu: hold userq kref during suspend and resume Zhu Lingshan
2026-08-28 9:53 ` [PATCH 09/10] drm/amdgpu: free userq by kref_put when fails to create Zhu Lingshan
2026-08-28 9:53 ` [PATCH 10/10] drm/amdgpu: take queue kref in userq_create to avoid UAF Zhu Lingshan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=847308aa-e4cb-405f-9e9c-dda3e4d28eef@amd.com \
--to=christian.koenig@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Ray.Huang@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=felix.kuehling@amd.com \
--cc=lingshan.zhu@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox