AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Danilo Krummrich" <dakr@kernel.org>
To: "Matthew Brost" <matthew.brost@intel.com>
Cc: vitaly.prosyak@amd.com, amd-gfx@lists.freedesktop.org,
	dri-devel@lists.freedesktop.org,
	"Christian König" <christian.koenig@amd.com>,
	"Alex Deucher" <alexander.deucher@amd.com>,
	"Philipp Stanner" <phasta@kernel.org>
Subject: Re: [PATCH 1/2] drm/sched: keep the current runqueue when no scheduler is ready
Date: Mon, 28 Sep 2026 10:52:44 +0200	[thread overview]
Message-ID: <DLQTLWC6YFPE.1FN6GSV00AWXM@kernel.org> (raw)
In-Reply-To: <arnPBVP3cnDNwTMn@gsse-cloud1.jf.intel.com>

On Mon Sep 28, 2026 at 4:20 AM CEST, Matthew Brost wrote:
> To me, this looks like a lifetime issue that should be fixed in AMDGPU.
> Either that, or we need to rework DRM to have proper lifetime management,
> or, of course, just deprecate drm_sched.

Agreed. The lifetime and ownership model in drm_sched is not fundamentally
wrong, but the implementation is not consistently keeping it up and drivers also
don't always honor it.

> In other words, we need refcounting so that drm_sched_fini() cannot be
> called while jobs are still in flight, nor can jobs be submitted before
> drm_sched_init(). Alternatively, drm_dep could serve as a replacement.

I'm not a huge fan of refcounting for those kind of things because it
fundamentally incentivises the wrong lifetime model in the context of the driver
model. The driver model requires a bounded lifetime scope, but refcounting
incentivises an unbounded lifetime model, which leads to other problems.

So, especially for the sake of deferring drm_sched_fini() from running jobs,
drivers still have to make sure that all jobs are torn down and drm_sched_fini()
is called *before* driver unbind completes. IOW, drivers should tear down the
hardware and hence all jobs latest in remove() and then call drm_sched_fini()
subsequently.

Refcounting does not provide a lot of value in this regard, because we have to
somehow guarantee that the hardware and all jobs are torn down at this specific
boundary anyway.

This is also my biggest concern about drm_dep, it seems to be designed with
exactly the idea of an unbounded lifetime model, which is not the correct design
for anything that represents a device resource (e.g. a GPU job).

Now, to be fair, refcounting is really the only mechanism that we have in C to
manage lifetimes, everything else is more or less just a convention. That said,
I'm not all against refcounting in general, but we have to be careful about how
it influences the design in terms of a bounded and unbounded lifetime model.

  reply	other threads:[~2026-09-28  8:52 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  1:59 [PATCH 1/2] drm/sched: keep the current runqueue when no scheduler is ready vitaly.prosyak
2026-09-28  2:20 ` Matthew Brost
2026-09-28  8:52   ` Danilo Krummrich [this message]
2026-09-28  7:49 ` Philipp Stanner
2026-09-28 10:12 ` Christian König

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLQTLWC6YFPE.1FN6GSV00AWXM@kernel.org \
    --to=dakr@kernel.org \
    --cc=alexander.deucher@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=christian.koenig@amd.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=phasta@kernel.org \
    --cc=vitaly.prosyak@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox