AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Christian König" <christian.koenig@amd.com>
To: Zhu Lingshan <lingshan.zhu@amd.com>,
	Alexander.Deucher@amd.com, felix.kuehling@amd.com
Cc: Ray.Huang@amd.com, amd-gfx@lists.freedesktop.org
Subject: Re: [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues
Date: Fri, 28 Aug 2026 15:09:38 +0200	[thread overview]
Message-ID: <18014f3c-5e0f-4ed7-b278-590cf60dfe03@amd.com> (raw)
In-Reply-To: <20260828095349.9797-3-lingshan.zhu@amd.com>

On 8/28/26 11:53, Zhu Lingshan wrote:
> The life cycle of a user queue is managed by its
> kref. However when destroy a userq manager,
> the kref_put of its queues in amdgpu_userq_mgr_fini
> may not be the last put, therefore the queues
> could be still alive after the userq manager
> has been destroyed, resulting in
> userq->userq_mgr use-after-free issues.
> 
> This commit fixes this problem by introduce a new
> counter refs representing for the number of its queues,
> and only free the userq_manager when refs == 0

Clear NAK to that one as well, this is just nonsense.

Christian.

> 
> Signed-off-by: Zhu Lingshan <lingshan.zhu@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 30 +++++++++++++++++++++++
>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h |  9 +++++++
>  2 files changed, 39 insertions(+)
> 
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> index e0639f844a8e..f398986a61a5 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> @@ -27,6 +27,7 @@
>  #include <linux/pm_runtime.h>
>  #include <linux/overflow.h>
>  #include <drm/drm_drv.h>
> +#include <linux/wait_bit.h>
>  
>  #include "amdgpu.h"
>  #include "amdgpu_reset.h"
> @@ -533,6 +534,17 @@ amdgpu_userq_get_doorbell_index(struct amdgpu_userq_mgr *uq_mgr,
>  	return r;
>  }
>  
> +static void amdgpu_userq_mgr_inc_refs(struct amdgpu_userq_mgr *uq_mgr)
> +{
> +	atomic_inc(&uq_mgr->refs);
> +}
> +
> +static void amdgpu_userq_mgr_dec_refs(struct amdgpu_userq_mgr *uq_mgr)
> +{
> +	if (atomic_dec_and_test(&uq_mgr->refs))
> +		wake_up_var(&uq_mgr->refs);
> +}
> +
>  static int
>  amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_queue *queue)
>  {
> @@ -594,6 +606,8 @@ static void amdgpu_userq_kref_destroy(struct kref *kref)
>  	r = amdgpu_userq_destroy(uq_mgr, queue);
>  	if (r)
>  		drm_file_err(uq_mgr->file, "Failed to destroy usermode queue %d\n", r);
> +
> +	amdgpu_userq_mgr_dec_refs(uq_mgr);
>  }
>  
>  struct amdgpu_usermode_queue *amdgpu_userq_get(struct amdgpu_userq_mgr *uq_mgr, u32 qid)
> @@ -707,6 +721,7 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
>  	queue->xcp_id = (fpriv->xcp_id != AMDGPU_XCP_NO_PARTITION) ?
>  				fpriv->xcp_id : 0;
>  	queue->userq_mgr = uq_mgr;
> +	amdgpu_userq_mgr_inc_refs(uq_mgr);
>  	INIT_DELAYED_WORK(&queue->hang_detect_work,
>  			  amdgpu_userq_hang_detect_work);
>  
> @@ -819,6 +834,7 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
>  free_queue:
>  	trace_amdgpu_userq_create_end(queue, r);
>  	kfree(queue);
> +	amdgpu_userq_mgr_dec_refs(uq_mgr);
>  err_pm_runtime:
>  	pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
>  	return r;
> @@ -1331,6 +1347,7 @@ int amdgpu_userq_mgr_init(struct amdgpu_userq_mgr *userq_mgr, struct drm_file *f
>  {
>  	mutex_init(&userq_mgr->userq_mutex);
>  	xa_init_flags(&userq_mgr->userq_xa, XA_FLAGS_ALLOC);
> +	atomic_set(&userq_mgr->refs, 0);
>  	userq_mgr->adev = adev;
>  	userq_mgr->file = file_priv;
>  	userq_mgr->proc_ctx_allocated = false;
> @@ -1380,6 +1397,19 @@ void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr)
>  		amdgpu_userq_put(queue);
>  	}
>  
> +	/*
> +	 * The above amdgpu_userq_put() may not be the last put
> +	 * of the kref of a user queue, therefore there could
> +	 * be some queues still alive even when the userq manager
> +	 * has been destroyed. This wait_evet() blocks
> +	 * amdgpu_userq_mgr_fini(), so keep userq_mgr alive
> +	 * while any queues holding it.
> +	 *
> +	 * This prevents queue->userq_mgr use-after-free issues.
> +	 */
> +	wait_var_event(&userq_mgr->refs,
> +		       !atomic_read_acquire(&userq_mgr->refs));
> +
>  	xa_destroy(&userq_mgr->userq_xa);
>  
>  	/*
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
> index 8fc73862f64e..a13d8d4dd5c7 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
> @@ -126,6 +126,15 @@ struct amdgpu_userq_mgr {
>  	 */
>  	struct xarray			userq_xa;
>  	struct mutex			userq_mutex;
> +
> +	/**
> +	 * @refs:
> +	 *
> +	 * Each queue increases this counter when join this manager,
> +	 * and decreases it when leave this manager.
> +	 */
> +	atomic_t			refs;
> +
>  	struct amdgpu_device		*adev;
>  	struct delayed_work		resume_work;
>  	struct drm_file			*file;


  reply	other threads:[~2026-08-28 13:09 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28  9:53 [PATCH 00/10] drm/amdgpu: secure userq lifecycle by its kref Zhu Lingshan
2026-08-28  9:53 ` [PATCH 01/10] drm/amdgpu: introduce amdgpu_lookup_queue_by_doorbell Zhu Lingshan
2026-08-28 13:08   ` Christian König
2026-08-28 15:59     ` Zhu, Lingshan
2026-08-28  9:53 ` [PATCH 02/10] drm/amdgpu: keep the userq manager alive as long as its queues Zhu Lingshan
2026-08-28 13:09   ` Christian König [this message]
2026-08-28 15:59     ` Zhu, Lingshan
2026-08-28 16:26       ` Christian König
2026-08-28  9:53 ` [PATCH 03/10] drm/amdgpu/gfx11: hold userq refs in private fault worker Zhu Lingshan
2026-08-28 13:11   ` Christian König
2026-08-28 15:59     ` Zhu, Lingshan
2026-08-28  9:53 ` [PATCH 04/10] drm/amdgpu/gfx12: " Zhu Lingshan
2026-08-28  9:53 ` [PATCH 05/10] drm/amdgpu: implement asynchronous userq destruction routine Zhu Lingshan
2026-08-28  9:53 ` [PATCH 06/10] drm/amdgpu: hold userq kref in MES reset Zhu Lingshan
2026-08-28  9:53 ` [PATCH 07/10] drm/amdgpu: hold userq kref during isolation scheduling Zhu Lingshan
2026-08-28  9:53 ` [PATCH 08/10] drm/amdgpu: hold userq kref during suspend and resume Zhu Lingshan
2026-08-28  9:53 ` [PATCH 09/10] drm/amdgpu: free userq by kref_put when fails to create Zhu Lingshan
2026-08-28  9:53 ` [PATCH 10/10] drm/amdgpu: take queue kref in userq_create to avoid UAF Zhu Lingshan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=18014f3c-5e0f-4ed7-b278-590cf60dfe03@amd.com \
    --to=christian.koenig@amd.com \
    --cc=Alexander.Deucher@amd.com \
    --cc=Ray.Huang@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=felix.kuehling@amd.com \
    --cc=lingshan.zhu@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox