AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Christian König" <christian.koenig@amd.com>
To: "Zhu, Lingshan" <lingshan.zhu@amd.com>,
	Alexander.Deucher@amd.com, felix.kuehling@amd.com
Cc: Ray.Huang@amd.com, amd-gfx@lists.freedesktop.org
Subject: Re: [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues
Date: Fri, 28 Aug 2026 18:26:06 +0200	[thread overview]
Message-ID: <847308aa-e4cb-405f-9e9c-dda3e4d28eef@amd.com> (raw)
In-Reply-To: <7f53ac18-ecb5-4eaa-8692-67b9237447b2@amd.com>

On 8/28/26 17:59, Zhu, Lingshan wrote:
> On 8/28/2026 9:09 PM, Christian König wrote:
> 
>> On 8/28/26 11:53, Zhu Lingshan wrote:
>>> The life cycle of a user queue is managed by its
>>> kref. However when destroy a userq manager,
>>> the kref_put of its queues in amdgpu_userq_mgr_fini
>>> may not be the last put, therefore the queues
>>> could be still alive after the userq manager
>>> has been destroyed, resulting in
>>> userq->userq_mgr use-after-free issues.
>>>
>>> This commit fixes this problem by introduce a new
>>> counter refs representing for the number of its queues,
>>> and only free the userq_manager when refs == 0
>> Clear NAK to that one as well, this is just nonsense.
> 
> It could be better to have some explanations.
> 
> I am not sure how to guarantee the put_kref in amdgpu_userq_mgr_fini
> is the last put and result in kref == 0, if not the last one,
> there can be userq->userq_mgr UAF bugs.

The rules are actually pretty simple:

The reference is for keeping the userq alive while IOCTLs happen. And IOCTL can only happen while the file and therefor the fpriv, userq_mgr etc... are still alive.

What can potentially be is that we also need to grab a reference from a work item, but in this case the fpriv/userq_mgr cleanup functions just need to cancel and wait for the work to finish.

There should *never* be a reference grabbed from interrupt context, explicitely because releasing that reference is also not possible from interrupt context. Instead xa_lock_irqsave() needs to be used to make sure that the userq stays alive while the interrupt processing happens.

Regards,
Christian.

> 
> Thanks
> Lingshan
> 
>> Christian.
>>
>>> Signed-off-by: Zhu Lingshan <lingshan.zhu@amd.com>
>>> ---
>>>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 30 +++++++++++++++++++++++
>>>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h |  9 +++++++
>>>  2 files changed, 39 insertions(+)
>>>
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>> index e0639f844a8e..f398986a61a5 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>> @@ -27,6 +27,7 @@
>>>  #include <linux/pm_runtime.h>
>>>  #include <linux/overflow.h>
>>>  #include <drm/drm_drv.h>
>>> +#include <linux/wait_bit.h>
>>>  
>>>  #include "amdgpu.h"
>>>  #include "amdgpu_reset.h"
>>> @@ -533,6 +534,17 @@ amdgpu_userq_get_doorbell_index(struct amdgpu_userq_mgr *uq_mgr,
>>>  	return r;
>>>  }
>>>  
>>> +static void amdgpu_userq_mgr_inc_refs(struct amdgpu_userq_mgr *uq_mgr)
>>> +{
>>> +	atomic_inc(&uq_mgr->refs);
>>> +}
>>> +
>>> +static void amdgpu_userq_mgr_dec_refs(struct amdgpu_userq_mgr *uq_mgr)
>>> +{
>>> +	if (atomic_dec_and_test(&uq_mgr->refs))
>>> +		wake_up_var(&uq_mgr->refs);
>>> +}
>>> +
>>>  static int
>>>  amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_queue *queue)
>>>  {
>>> @@ -594,6 +606,8 @@ static void amdgpu_userq_kref_destroy(struct kref *kref)
>>>  	r = amdgpu_userq_destroy(uq_mgr, queue);
>>>  	if (r)
>>>  		drm_file_err(uq_mgr->file, "Failed to destroy usermode queue %d\n", r);
>>> +
>>> +	amdgpu_userq_mgr_dec_refs(uq_mgr);
>>>  }
>>>  
>>>  struct amdgpu_usermode_queue *amdgpu_userq_get(struct amdgpu_userq_mgr *uq_mgr, u32 qid)
>>> @@ -707,6 +721,7 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
>>>  	queue->xcp_id = (fpriv->xcp_id != AMDGPU_XCP_NO_PARTITION) ?
>>>  				fpriv->xcp_id : 0;
>>>  	queue->userq_mgr = uq_mgr;
>>> +	amdgpu_userq_mgr_inc_refs(uq_mgr);
>>>  	INIT_DELAYED_WORK(&queue->hang_detect_work,
>>>  			  amdgpu_userq_hang_detect_work);
>>>  
>>> @@ -819,6 +834,7 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
>>>  free_queue:
>>>  	trace_amdgpu_userq_create_end(queue, r);
>>>  	kfree(queue);
>>> +	amdgpu_userq_mgr_dec_refs(uq_mgr);
>>>  err_pm_runtime:
>>>  	pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
>>>  	return r;
>>> @@ -1331,6 +1347,7 @@ int amdgpu_userq_mgr_init(struct amdgpu_userq_mgr *userq_mgr, struct drm_file *f
>>>  {
>>>  	mutex_init(&userq_mgr->userq_mutex);
>>>  	xa_init_flags(&userq_mgr->userq_xa, XA_FLAGS_ALLOC);
>>> +	atomic_set(&userq_mgr->refs, 0);
>>>  	userq_mgr->adev = adev;
>>>  	userq_mgr->file = file_priv;
>>>  	userq_mgr->proc_ctx_allocated = false;
>>> @@ -1380,6 +1397,19 @@ void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr)
>>>  		amdgpu_userq_put(queue);
>>>  	}
>>>  
>>> +	/*
>>> +	 * The above amdgpu_userq_put() may not be the last put
>>> +	 * of the kref of a user queue, therefore there could
>>> +	 * be some queues still alive even when the userq manager
>>> +	 * has been destroyed. This wait_evet() blocks
>>> +	 * amdgpu_userq_mgr_fini(), so keep userq_mgr alive
>>> +	 * while any queues holding it.
>>> +	 *
>>> +	 * This prevents queue->userq_mgr use-after-free issues.
>>> +	 */
>>> +	wait_var_event(&userq_mgr->refs,
>>> +		       !atomic_read_acquire(&userq_mgr->refs));
>>> +
>>>  	xa_destroy(&userq_mgr->userq_xa);
>>>  
>>>  	/*
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>> index 8fc73862f64e..a13d8d4dd5c7 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>> @@ -126,6 +126,15 @@ struct amdgpu_userq_mgr {
>>>  	 */
>>>  	struct xarray			userq_xa;
>>>  	struct mutex			userq_mutex;
>>> +
>>> +	/**
>>> +	 * @refs:
>>> +	 *
>>> +	 * Each queue increases this counter when join this manager,
>>> +	 * and decreases it when leave this manager.
>>> +	 */
>>> +	atomic_t			refs;
>>> +
>>>  	struct amdgpu_device		*adev;
>>>  	struct delayed_work		resume_work;
>>>  	struct drm_file			*file;


  reply	other threads:[~2026-08-28 16:26 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28  9:53 [PATCH 00/10] drm/amdgpu: secure userq lifecycle by its kref Zhu Lingshan
2026-08-28  9:53 ` [PATCH 01/10] drm/amdgpu: introduce amdgpu_lookup_queue_by_doorbell Zhu Lingshan
2026-08-28 13:08   ` Christian König
2026-08-28 15:59     ` Zhu, Lingshan
2026-08-28  9:53 ` [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues Zhu Lingshan
2026-08-28 13:09   ` Christian König
2026-08-28 15:59     ` Zhu, Lingshan
2026-08-28 16:26       ` Christian König [this message]
2026-08-28  9:53 ` [PATCH 03/10] drm/amdgpu/gfx11: hold userq refs in private fault worker Zhu Lingshan
2026-08-28 13:11   ` Christian König
2026-08-28 15:59     ` Zhu, Lingshan
2026-08-28  9:53 ` [PATCH 04/10] drm/amdgpu/gfx12: " Zhu Lingshan
2026-08-28  9:53 ` [PATCH 05/10] drm/amdgpu: implement asynchronous userq destruction routine Zhu Lingshan
2026-08-28  9:53 ` [PATCH 06/10] drm/amdgpu: hold userq kref in MES reset Zhu Lingshan
2026-08-28  9:53 ` [PATCH 07/10] drm/amdgpu: hold userq kref during isolation scheduling Zhu Lingshan
2026-08-28  9:53 ` [PATCH 08/10] drm/amdgpu: hold userq kref during suspend and resume Zhu Lingshan
2026-08-28  9:53 ` [PATCH 09/10] drm/amdgpu: free userq by kref_put when fails to create Zhu Lingshan
2026-08-28  9:53 ` [PATCH 10/10] drm/amdgpu: take queue kref in userq_create to avoid UAF Zhu Lingshan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=847308aa-e4cb-405f-9e9c-dda3e4d28eef@amd.com \
    --to=christian.koenig@amd.com \
    --cc=Alexander.Deucher@amd.com \
    --cc=Ray.Huang@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=felix.kuehling@amd.com \
    --cc=lingshan.zhu@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox