From: sashiko-bot@kernel.org
To: "Tvrtko Ursulin" <tvrtko.ursulin@igalia.com>
Cc: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org
Subject: Re: [PATCH 8/9] drm/sched: Replace completion with a flush
Date: Sat, 03 Oct 2026 01:33:13 +0000 [thread overview]
Message-ID: <20261003013314.507811F00893@smtp.kernel.org> (raw)
In-Reply-To: <20261002154713.77591-9-tvrtko.ursulin@igalia.com>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] drm/sched: cross-entity deadlock via flush_work() in drm_sched_entity_kill()
--
commit 3786a3d89362d00183a97f456e4eaeb1dff144f4
Author: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
drm/sched: Replace completion with a flush
This commit removes the per-entity completion used to synchronize entity
teardown, replacing it with a direct flush of the scheduler's run job
worker. The goal is to achieve the same synchronization guarantee without
requiring the scheduler to signal a completion when an entity goes idle.
> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
> --- a/drivers/gpu/drm/scheduler/sched_entity.c
> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
[ ... ]
> @@ -304,16 +299,23 @@ static void drm_sched_entity_kill_jobs_cb(struct dma_fence *f,
> void drm_sched_entity_kill(struct drm_sched_entity *entity)
> {
[ ... ]
> - /* Make sure this entity is not used by the scheduler at the moment */
> - wait_for_completion(&entity->entity_idle);
> + /*
> + * Make sure this entity is not used by the scheduler at the moment.
> + *
> + * Scheduler is guaranteed to be stable after the entity was stopped and
> + * removed from the run-queue.
> + */
> + if (sched)
> + drm_sched_flush_run_work(sched);
[Severity: High]
Could this introduce a cross-entity deadlock regression during context teardown?
Because drm_sched_flush_run_work calls flush_work on the scheduler's shared
worker, it synchronously waits for the worker to finish its current execution.
This might cause it to wait on a job belonging to a completely different
entity.
If a process destroys its device context and triggers drm_sched_entity_kill
while holding a driver-specific lock, and the shared worker is concurrently
running drm_sched_run_job_work for an unrelated entity, that driver-provided
run_job callback might attempt to acquire the same driver lock.
The original code used a per-entity completion (entity_idle) which explicitly
avoided waiting on other entities. Does this flush create a new cross-entity
dependency that could lead to a deadlock?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261002154713.77591-1-tvrtko.ursulin@igalia.com?part=8
prev parent reply other threads:[~2026-10-03 1:33 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <20261002154713.77591-1-tvrtko.ursulin@igalia.com>
2026-10-02 15:47 ` [PATCH 8/9] drm/sched: Replace completion with a flush Tvrtko Ursulin
2026-10-03 1:33 ` sashiko-bot [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261003013314.507811F00893@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=tvrtko.ursulin@igalia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox