Intel-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Belgaumkar, Vinay" <vinay.belgaumkar@intel.com>
To: Umesh Nerlige Ramappa <umesh.nerlige.ramappa@intel.com>,
	<intel-gfx@lists.freedesktop.org>
Subject: Re: [PATCH 1/2] drm/i915/guc: Use a delayed wakeref put in guc_engine_busyness()
Date: Tue, 22 Sep 2026 15:09:03 -0700	[thread overview]
Message-ID: <b0337ea8-58db-4cea-a355-54f09a2e2d06@intel.com> (raw)
In-Reply-To: <20260910184253.1231313-5-umesh.nerlige.ramappa@intel.com>


On 9/10/2026 11:42 AM, Umesh Nerlige Ramappa wrote:
> perf can invoke the PMU callbacks from the scheduler with the runqueue
> lock held, e.g. pmu->del() via perf_cgroup_switch() from
> finish_task_switch(). guc_engine_busyness() drops its GT wakeref there
> with intel_gt_pm_put_async(), which lands in mod_delayed_work() with a
> zero delay. That queues the work immediately and wakes a worker via
> try_to_wake_up(), which then deadlocks on the runqueue lock we are
> already holding. The put also happens under guc->timestamp.lock, so the
> stall blocks any concurrent __guc_context_update_stats() too, and the
> machine hard locks up.
>
> Only reachable when a concurrent put drops the refcount to 1 inside the
> get_if_awake()/put_async() window, which makes it rare and load
> dependent.
>
> Add intel_gt_pm_put_delay() and use it with a 1 jiffy delay. A non-zero
> delay takes the add_timer_on() path in __queue_delayed_work() instead,
> so no task is ever woken from here.
>
> Fixes: 77cdd054dd2c ("drm/i915/pmu: Connect engine busyness stats from GuC to pmu")
> Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/issues/16984
> Signed-off-by: Umesh Nerlige Ramappa <umesh.nerlige.ramappa@intel.com>
> Assisted-by: Claude:claude-opus-5
> ---
>   drivers/gpu/drm/i915/gt/intel_gt_pm.h         | 24 +++++++++++++++++++
>   .../gpu/drm/i915/gt/uc/intel_guc_submission.c |  7 +++++-
>   2 files changed, 30 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/gpu/drm/i915/gt/intel_gt_pm.h b/drivers/gpu/drm/i915/gt/intel_gt_pm.h
> index 6f25c747bc29..24c8b014864d 100644
> --- a/drivers/gpu/drm/i915/gt/intel_gt_pm.h
> +++ b/drivers/gpu/drm/i915/gt/intel_gt_pm.h
> @@ -72,6 +72,30 @@ static inline void intel_gt_pm_put_async(struct intel_gt *gt, intel_wakeref_t ha
>   	intel_gt_pm_put_async_untracked(gt);
>   }
>   
> +/**
> + * intel_gt_pm_put_delay - release the GT wakeref, deferred by a timer
> + *
> + * @gt: pointer to the gt
> + * @handle: the handle returned by the matching intel_gt_pm_get*()
> + * @delay: delay, in jiffies, before the release is processed
> + *
> + * As intel_gt_pm_put_async(), except that dropping the last reference arms a
> + * timer instead of queueing the release immediately.
> + *
> + * intel_gt_pm_put_async() ends up in mod_delayed_work() with a zero delay,
> + * which queues the work straight away and so wakes a workqueue worker through
> + * try_to_wake_up(). That deadlocks if the caller already holds a runqueue
> + * lock, since try_to_wake_up() then tries to take it again. Callers which can
> + * run from inside the scheduler must use this variant with a non-zero @delay,
> + * which takes the add_timer_on() path and never wakes a task.
> + */
> +static inline void intel_gt_pm_put_delay(struct intel_gt *gt, intel_wakeref_t handle,
> +					 unsigned long delay)

Would a 0 delay make it the same as async? Should we have just one 
function based on delay value then?

Thanks,

Vinay.

> +{
> +	intel_wakeref_untrack(&gt->wakeref, handle);
> +	intel_wakeref_put_delay(&gt->wakeref, delay);
> +}
> +
>   #define with_intel_gt_pm(gt, wf) \
>   	for ((wf) = intel_gt_pm_get(gt); (wf); intel_gt_pm_put((gt), (wf)), (wf) = NULL)
>   
> diff --git a/drivers/gpu/drm/i915/gt/uc/intel_guc_submission.c b/drivers/gpu/drm/i915/gt/uc/intel_guc_submission.c
> index 788e59cdfac9..32d722c378e8 100644
> --- a/drivers/gpu/drm/i915/gt/uc/intel_guc_submission.c
> +++ b/drivers/gpu/drm/i915/gt/uc/intel_guc_submission.c
> @@ -1361,7 +1361,12 @@ static ktime_t guc_engine_busyness(struct intel_engine_cs *engine, ktime_t *now)
>   		 */
>   		guc_update_engine_gt_clks(engine);
>   		guc_update_pm_timestamp(guc, now);
> -		intel_gt_pm_put_async(gt, wakeref);
> +		/*
> +		 * We are reached from the perf callbacks, which perf may invoke
> +		 * from the scheduler with the runqueue lock held. Defer the put
> +		 * on a timer so that it can never wake a task from here.
> +		 */
> +		intel_gt_pm_put_delay(gt, wakeref, 1);
>   		if (i915_reset_count(gpu_error) != reset_count) {
>   			*stats = stats_saved;
>   			guc->timestamp.gt_stamp = gt_stamp_saved;

  reply	other threads:[~2026-09-22 22:09 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 18:42 [PATCH 0/2] Fixup deadlock scenarios in the PMU path Umesh Nerlige Ramappa
2026-09-10 18:42 ` [PATCH 1/2] drm/i915/guc: Use a delayed wakeref put in guc_engine_busyness() Umesh Nerlige Ramappa
2026-09-22 22:09   ` Belgaumkar, Vinay [this message]
2026-09-24 17:55     ` Umesh Nerlige Ramappa
2026-09-10 18:42 ` [PATCH 2/2] drm/i915/pmu: Use a delayed wakeref put in get_rc6() Umesh Nerlige Ramappa
2026-09-10 18:54   ` sashiko-bot
2026-09-16 21:50     ` Umesh Nerlige Ramappa
2026-09-22 22:11   ` Belgaumkar, Vinay
2026-09-24 17:02     ` Umesh Nerlige Ramappa
2026-09-10 22:13 ` ✓ i915.CI.BAT: success for Fixup deadlock scenarios in the PMU path Patchwork
2026-09-11 17:04 ` ✗ i915.CI.Full: failure " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=b0337ea8-58db-4cea-a355-54f09a2e2d06@intel.com \
    --to=vinay.belgaumkar@intel.com \
    --cc=intel-gfx@lists.freedesktop.org \
    --cc=umesh.nerlige.ramappa@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox