Intel-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Daniel Vetter <daniel@ffwll.ch>
To: Chris Wilson <chris@chris-wilson.co.uk>
Cc: intel-gfx@lists.freedesktop.org, stable@vger.kernel.org,
	Mika Kuoppala <mika.kuoppala@intel.com>
Subject: Re: [PATCH] drm/i915: Fix erroneous dereference of batch_obj inside reset_status
Date: Wed, 4 Dec 2013 13:18:42 +0100	[thread overview]
Message-ID: <20131204121842.GB27344@phenom.ffwll.local> (raw)
In-Reply-To: <1386157029-5954-1-git-send-email-chris@chris-wilson.co.uk>

On Wed, Dec 04, 2013 at 11:37:09AM +0000, Chris Wilson wrote:
> As the rings may be processed and their requests deallocated in a
> different order to the natural retirement during a reset,
> 
> /* Whilst this request exists, batch_obj will be on the
>  * active_list, and so will hold the active reference. Only when this
>  * request is retired will the the batch_obj be moved onto the
>  * inactive_list and lose its active reference. Hence we do not need
>  * to explicitly hold another reference here.
>  */
> 
> is violated, and the batch_obj may be dereferenced after it had been
> freed on another ring. This can be simply avoided by processing the
> status update prior to deallocating any requests.
> 
> Fixes regression (a possible OOPS following a GPU hang) from
> commit aa60c664e6df502578454621c3a9b1f087ff8d25
> Author: Mika Kuoppala <mika.kuoppala@linux.intel.com>
> Date:   Wed Jun 12 15:13:20 2013 +0300
> 
>     drm/i915: find guilty batch buffer on ring resets
> 
> Signed-off-by: Chris Wilson <chris@chris-wilson.co.uk>
> Cc: Mika Kuoppala <mika.kuoppala@intel.com>
> Cc: stable@vger.kernel.org
> ---
>  drivers/gpu/drm/i915/i915_gem.c | 29 +++++++++++++++++++----------
>  1 file changed, 19 insertions(+), 10 deletions(-)
> 
> diff --git a/drivers/gpu/drm/i915/i915_gem.c b/drivers/gpu/drm/i915/i915_gem.c
> index ec4502034203..c1e481d36575 100644
> --- a/drivers/gpu/drm/i915/i915_gem.c
> +++ b/drivers/gpu/drm/i915/i915_gem.c
> @@ -2442,15 +2442,24 @@ static void i915_gem_free_request(struct drm_i915_gem_request *request)
>  	kfree(request);
>  }
>  
> -static void i915_gem_reset_ring_lists(struct drm_i915_private *dev_priv,
> -				      struct intel_ring_buffer *ring)
> +static void i915_gem_reset_ring_status(struct drm_i915_private *dev_priv,
> +				       struct intel_ring_buffer *ring)
>  {
> -	u32 completed_seqno;
> -	u32 acthd;
> +	u32 completed_seqno = ring->get_seqno(ring, false);
> +	u32 acthd = intel_ring_get_active_head(ring);
> +	struct drm_i915_gem_request *request;
> +
> +	list_for_each_entry(request, &ring->request_list, list) {
> +		if (i915_seqno_passed(completed_seqno, request->seqno))
> +			continue;
>  
> -	acthd = intel_ring_get_active_head(ring);
> -	completed_seqno = ring->get_seqno(ring, false);
> +		i915_set_reset_status(ring, request, acthd);
> +	}
> +}

Indeed the fix in the gem reset code is a bit simpler than what I've
feared. We still have fairly tricky code which depends upon that implicit
reference in non-obvious ways. So I still think Mika's refcount patch with
the comments updated is the better approach.
-Daniel


>  
> +static void i915_gem_reset_ring_cleanup(struct drm_i915_private *dev_priv,
> +					struct intel_ring_buffer *ring)
> +{
>  	while (!list_empty(&ring->request_list)) {
>  		struct drm_i915_gem_request *request;
>  
> @@ -2458,9 +2467,6 @@ static void i915_gem_reset_ring_lists(struct drm_i915_private *dev_priv,
>  					   struct drm_i915_gem_request,
>  					   list);
>  
> -		if (request->seqno > completed_seqno)
> -			i915_set_reset_status(ring, request, acthd);
> -
>  		i915_gem_free_request(request);
>  	}
>  
> @@ -2503,7 +2509,10 @@ void i915_gem_reset(struct drm_device *dev)
>  	int i;
>  
>  	for_each_ring(ring, dev_priv, i)
> -		i915_gem_reset_ring_lists(dev_priv, ring);
> +		i915_gem_reset_ring_status(dev_priv, ring);
> +
> +	for_each_ring(ring, dev_priv, i)
> +		i915_gem_reset_ring_cleanup(dev_priv, ring);
>  
>  	i915_gem_cleanup_ringbuffer(dev);
>  
> -- 
> 1.8.5.1
> 
> _______________________________________________
> Intel-gfx mailing list
> Intel-gfx@lists.freedesktop.org
> http://lists.freedesktop.org/mailman/listinfo/intel-gfx

-- 
Daniel Vetter
Software Engineer, Intel Corporation
+41 (0) 79 365 57 48 - http://blog.ffwll.ch

  reply	other threads:[~2013-12-04 12:17 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2013-12-04 11:37 [PATCH] drm/i915: Fix erroneous dereference of batch_obj inside reset_status Chris Wilson
2013-12-04 12:18 ` Daniel Vetter [this message]
2013-12-04 12:32   ` [Intel-gfx] " Chris Wilson
2013-12-05 16:07 ` Mika Kuoppala
2013-12-05 16:22   ` Chris Wilson
2013-12-12  9:54     ` Daniel Vetter

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20131204121842.GB27344@phenom.ffwll.local \
    --to=daniel@ffwll.ch \
    --cc=chris@chris-wilson.co.uk \
    --cc=intel-gfx@lists.freedesktop.org \
    --cc=mika.kuoppala@intel.com \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox