amd-gfx.lists.freedesktop.org archive mirror
 help / color / mirror / Atom feed
* [PATCH 1/2] drm/sched: keep the current runqueue when no scheduler is ready
@ 2026-09-28  1:59 vitaly.prosyak
  2026-09-28  2:20 ` Matthew Brost
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: vitaly.prosyak @ 2026-09-28  1:59 UTC (permalink / raw)
  To: amd-gfx, dri-devel
  Cc: Vitaly Prosyak, Christian König, Alex Deucher, Matthew Brost,
	Danilo Krummrich, Philipp Stanner

From: Vitaly Prosyak <vitaly.prosyak@amd.com>

The IGT amd_dispatch test exposed a NULL pointer dereference in the
AMDGPU CS submission path when the GPU schedulers were not ready.

drm_sched_pick_best() returns NULL when every scheduler in an entity's
list is marked not ready. drm_sched_entity_select_rq() then replaces
the entity's existing runqueue with NULL.

A subsequent drm_sched_job_arm() retains that invalid runqueue.
When AMDGPU CS submission calls drm_sched_entity_push_job(), the
scheduler pointer derived from entity->rq is invalid and the access
to sched->score faults. The reported oops shows the sequence:

    [drm] scheduler comp_1.1.0 is not ready, skipping
    [drm] scheduler comp_1.2.0 is not ready, skipping
    BUG: kernel NULL pointer dereference, address: 0000000000000268
    RIP: drm_sched_entity_push_job+0x4f/0x2b0 [gpu_sched]
    Call Trace:
      amdgpu_cs_ioctl+0x1e9e/0x2530 [amdgpu]

Keep the previously selected runqueue when no ready replacement is
found. This prevents scheduler selection from turning a valid entity
runqueue into NULL; it does not make a stopped scheduler ready or
guarantee that the submitted job will execute.

Cc: Christian König <christian.koenig@amd.com>
Cc: Alex Deucher <alexander.deucher@amd.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Danilo Krummrich <dakr@kernel.org>
Cc: Philipp Stanner <phasta@kernel.org>
Signed-off-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
---
 drivers/gpu/drm/scheduler/sched_entity.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
index 4ebb513255ed..b11e1dddabd0 100644
--- a/drivers/gpu/drm/scheduler/sched_entity.c
+++ b/drivers/gpu/drm/scheduler/sched_entity.c
@@ -584,8 +584,8 @@ void drm_sched_entity_select_rq(struct drm_sched_entity *entity)
 
 	spin_lock(&entity->lock);
 	sched = drm_sched_pick_best(entity->sched_list, entity->num_sched_list);
-	rq = sched ? &sched->rq : NULL;
-	if (rq != entity->rq) {
+	if (sched && &sched->rq != entity->rq) {
+		rq = &sched->rq;
 		drm_sched_rq_remove_entity(entity->rq, entity);
 		entity->rq = rq;
 	}
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-28 10:12 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28  1:59 [PATCH 1/2] drm/sched: keep the current runqueue when no scheduler is ready vitaly.prosyak
2026-09-28  2:20 ` Matthew Brost
2026-09-28  8:52   ` Danilo Krummrich
2026-09-28  7:49 ` Philipp Stanner
2026-09-28 10:12 ` Christian König

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).