From: Matthew Brost <matthew.brost@intel.com>
To: <vitaly.prosyak@amd.com>
Cc: amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org,
"Christian König" <christian.koenig@amd.com>,
"Alex Deucher" <alexander.deucher@amd.com>,
"Danilo Krummrich" <dakr@kernel.org>,
"Philipp Stanner" <phasta@kernel.org>
Subject: Re: [PATCH 1/2] drm/sched: keep the current runqueue when no scheduler is ready
Date: Sun, 27 Sep 2026 19:20:53 -0700 [thread overview]
Message-ID: <arnPBVP3cnDNwTMn@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260928020018.120503-1-vitaly.prosyak@amd.com>
On Sun, Sep 27, 2026 at 09:59:02PM -0400, vitaly.prosyak@amd.com wrote:
> From: Vitaly Prosyak <vitaly.prosyak@amd.com>
>
> The IGT amd_dispatch test exposed a NULL pointer dereference in the
> AMDGPU CS submission path when the GPU schedulers were not ready.
>
> drm_sched_pick_best() returns NULL when every scheduler in an entity's
> list is marked not ready. drm_sched_entity_select_rq() then replaces
> the entity's existing runqueue with NULL.
>
> A subsequent drm_sched_job_arm() retains that invalid runqueue.
> When AMDGPU CS submission calls drm_sched_entity_push_job(), the
> scheduler pointer derived from entity->rq is invalid and the access
> to sched->score faults. The reported oops shows the sequence:
>
> [drm] scheduler comp_1.1.0 is not ready, skipping
> [drm] scheduler comp_1.2.0 is not ready, skipping
> BUG: kernel NULL pointer dereference, address: 0000000000000268
> RIP: drm_sched_entity_push_job+0x4f/0x2b0 [gpu_sched]
> Call Trace:
> amdgpu_cs_ioctl+0x1e9e/0x2530 [amdgpu]
>
> Keep the previously selected runqueue when no ready replacement is
> found. This prevents scheduler selection from turning a valid entity
> runqueue into NULL; it does not make a stopped scheduler ready or
> guarantee that the submitted job will execute.
>
> Cc: Christian König <christian.koenig@amd.com>
> Cc: Alex Deucher <alexander.deucher@amd.com>
> Cc: Matthew Brost <matthew.brost@intel.com>
> Cc: Danilo Krummrich <dakr@kernel.org>
> Cc: Philipp Stanner <phasta@kernel.org>
> Signed-off-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
> ---
> drivers/gpu/drm/scheduler/sched_entity.c | 4 ++--
> 1 file changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
> index 4ebb513255ed..b11e1dddabd0 100644
> --- a/drivers/gpu/drm/scheduler/sched_entity.c
> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
> @@ -584,8 +584,8 @@ void drm_sched_entity_select_rq(struct drm_sched_entity *entity)
>
> spin_lock(&entity->lock);
> sched = drm_sched_pick_best(entity->sched_list, entity->num_sched_list);
This entire thing is broken, badly.
So I assume this pops too in drm_sched_pick_best?
969 if (!sched->ready) {
970 DRM_WARN("scheduler %s is not ready, skipping",
971 sched->name);
972 continue;
973 }
To me, this looks like a lifetime issue that should be fixed in AMDGPU.
Either that, or we need to rework DRM to have proper lifetime management,
or, of course, just deprecate drm_sched.
In other words, we need refcounting so that drm_sched_fini() cannot be
called while jobs are still in flight, nor can jobs be submitted before
drm_sched_init(). Alternatively, drm_dep could serve as a replacement.
So this is a NAK from me. As it stands, this is papering over a larger
issue. You won't hit a NULL pointer dereference, but sched->ready on
entity->rq will be false. At that point, what does sched->ready even
mean?
Matt
> - rq = sched ? &sched->rq : NULL;
> - if (rq != entity->rq) {
> + if (sched && &sched->rq != entity->rq) {
> + rq = &sched->rq;
> drm_sched_rq_remove_entity(entity->rq, entity);
> entity->rq = rq;
> }
> --
> 2.43.0
>
next prev parent reply other threads:[~2026-09-28 2:21 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 1:59 [PATCH 1/2] drm/sched: keep the current runqueue when no scheduler is ready vitaly.prosyak
2026-09-28 2:20 ` Matthew Brost [this message]
2026-09-28 8:52 ` Danilo Krummrich
2026-09-28 7:49 ` Philipp Stanner
2026-09-28 10:12 ` Christian König
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arnPBVP3cnDNwTMn@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=alexander.deucher@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
--cc=dakr@kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=phasta@kernel.org \
--cc=vitaly.prosyak@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox