From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 858913B14A3; Wed, 23 Sep 2026 14:35:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790174109; cv=none; b=VjH4wGIMWwzE8HwJIqZ1kAIYSUQs3fhsGfUnMQy9+nttQcVJkbTNb2M8TAZHqY34/wkphmLprg8rgmwR+uuqbfYjthpSs50o3djNGCcbkoy737O3937HghA4QXYmVTIwYVTqmHLrh3CKCeSPIpGP5LMqy2LmAZ8daw4UWPuMA0U= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790174109; c=relaxed/simple; bh=Cka5WVyO8V8J+OC+pM6zXBfNMYLSC2EI6xRWy8kTS4I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=G1Rigi59ovvVRG3DRkzufHZiwfnP1Mk5FANCVq8E8mV3/JaIhsNcjkH3AoxHSB8CG5cIKe6yVtKmqxerFcLYCNs2udd7bFPkkzGo9eXaXXOoeS4IxmxSSaZ6URxpbroZIJXP2Tz0SPa5ibHVh9pfKyG30jssuYk5q4ZjwC5IE58= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b=z9LtGx5C; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b="z9LtGx5C" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DA3FF1F000FF; Wed, 23 Sep 2026 14:35:07 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linuxfoundation.org; s=korg; t=1790174108; bh=2fUG7XhlPchssdnHjQWsGh2yuvCLGvX1hIxiKKJGB4Q=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=z9LtGx5CXAtcl4XUr9z0GPu/2rwDEc3r0/n7TZgl+glchfCKFh8Mf+RGIsA1jdK+0 Qpe1ZQxSiWbBuf6plWx799XTzdYq9jHS9KVxysnVDM137GqS/eFQMCEalR5ZRVicK3 j3feI1meCthvylgR1mlxzqbyETHmdX4hNcThm2gs= From: Greg Kroah-Hartman To: stable@vger.kernel.org Cc: Greg Kroah-Hartman , patches@lists.linux.dev, Tvrtko Ursulin , Luke.Wildhardt@proton.me, =?UTF-8?q?Christian=20K=C3=B6nig?= , Danilo Krummrich , Philipp Stanner , Pierre-Eric Pelloux-Prayer , Matthew Brost , Vitaly Prosyak Subject: [PATCH 7.2 406/438] drm/sched: Fix virtual runtime race Date: Wed, 23 Sep 2026 16:07:07 +0200 Message-ID: <20260923140655.420759009@linuxfoundation.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260923140644.756254324@linuxfoundation.org> References: <20260923140644.756254324@linuxfoundation.org> User-Agent: quilt/0.69 X-stable: review X-Patchwork-Hint: ignore Precedence: bulk X-Mailing-List: patches@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 7.2-stable review patch. If anyone has any objections, please let me know. ------------------ From: Tvrtko Ursulin commit 2ab510e63197360945f915dd5631a77c63ac6b27 upstream. Prevent pushing a new job to an entity seeing it being the first in the queue, and hence entering the drm_sched_rq_add_entity() path, if the pop side in drm_sched_entity_pop_job() has just de-queued the job but not yet updated the saved virtual time. Restoring the unsaved virtual time, which is at this point not a delta but still an absolute value, pushes the said entity to the rear of the run queue for a potentially very long time. We close this race by pulling the locked sections out to encompass both the queue push/pop and corresponding rbtree management. This is aligned with the future direction to replace the current lockless job queue with one of the fully locked standard list primitives. Signed-off-by: Tvrtko Ursulin Fixes: 2fa4d8e2c109 ("drm/sched: Add fair scheduling policy") Suggested-by: Luke.Wildhardt@proton.me # via Claude Opus Tested-by: Luke.Wildhardt@proton.me Cc: Christian König Cc: Danilo Krummrich Cc: Philipp Stanner Cc: Pierre-Eric Pelloux-Prayer Cc: Matthew Brost Cc: Vitaly Prosyak Cc: stable@vger.kernel.org # v7.2+ [phasta: commit title] Signed-off-by: Philipp Stanner Link: https://patch.msgid.link/20260915150557.62847-1-tvrtko.ursulin@igalia.com Signed-off-by: Greg Kroah-Hartman --- drivers/gpu/drm/scheduler/sched_entity.c | 8 +++++++- drivers/gpu/drm/scheduler/sched_rq.c | 20 +++++++++----------- 2 files changed, 16 insertions(+), 12 deletions(-) --- a/drivers/gpu/drm/scheduler/sched_entity.c +++ b/drivers/gpu/drm/scheduler/sched_entity.c @@ -566,9 +566,10 @@ struct drm_sched_job *drm_sched_entity_p */ smp_wmb(); + spin_lock(&entity->lock); spsc_queue_pop(&entity->job_queue); - drm_sched_rq_pop_entity(entity); + spin_unlock(&entity->lock); /* Jobs and entities might have different lifecycles. Since we're * removing the job from the entities queue, set the jobs entity pointer @@ -654,6 +655,9 @@ void drm_sched_entity_push_job(struct dr * Make sure to set the submit_ts first, to avoid a race. */ sched_job->submit_ts = submit_ts = ktime_get(); + + spin_lock(&entity->lock); + first = spsc_queue_push(&entity->job_queue, &sched_job->queue_node); /* first job wakes up scheduler */ @@ -664,5 +668,7 @@ void drm_sched_entity_push_job(struct dr if (sched) drm_sched_wakeup(sched); } + + spin_unlock(&entity->lock); } EXPORT_SYMBOL(drm_sched_entity_push_job); --- a/drivers/gpu/drm/scheduler/sched_rq.c +++ b/drivers/gpu/drm/scheduler/sched_rq.c @@ -257,19 +257,17 @@ static ktime_t drm_sched_entity_get_job_ struct drm_gpu_scheduler * drm_sched_rq_add_entity(struct drm_sched_entity *entity, ktime_t ts) { + struct drm_sched_rq *rq = entity->rq; struct drm_gpu_scheduler *sched; - struct drm_sched_rq *rq; /* Add the entity to the run queue */ - spin_lock(&entity->lock); - if (entity->stopped) { - spin_unlock(&entity->lock); + lockdep_assert_held(&entity->lock); + if (entity->stopped) { DRM_ERROR("Trying to push to a killed entity\n"); return NULL; } - rq = entity->rq; spin_lock(&rq->lock); sched = rq->sched; @@ -289,7 +287,6 @@ drm_sched_rq_add_entity(struct drm_sched drm_sched_rq_update_fifo_locked(entity, rq, ts); spin_unlock(&rq->lock); - spin_unlock(&entity->lock); return sched; } @@ -343,16 +340,17 @@ drm_sched_rq_next_rr_ts(struct drm_sched */ void drm_sched_rq_pop_entity(struct drm_sched_entity *entity) { + struct drm_sched_rq *rq = entity->rq; struct drm_sched_job *next_job; - struct drm_sched_rq *rq; + + lockdep_assert_held(&entity->lock); + + spin_lock(&rq->lock); /* * Update the entity's location in the min heap according to * the timestamp of the next job, if any. */ - spin_lock(&entity->lock); - rq = entity->rq; - spin_lock(&rq->lock); next_job = drm_sched_entity_queue_peek(entity); if (next_job) { ktime_t ts; @@ -375,8 +373,8 @@ void drm_sched_rq_pop_entity(struct drm_ drm_sched_entity_save_vruntime(entity, min_vruntime); } } + spin_unlock(&rq->lock); - spin_unlock(&entity->lock); } /**