From: Daniel Vetter <daniel@ffwll.ch>
To: Chris Wilson <chris@chris-wilson.co.uk>
Cc: Daniel Vetter <daniel.vetter@ffwll.ch>,
Intel Graphics Development <intel-gfx@lists.freedesktop.org>,
stable@vger.kernel.org
Subject: Re: [Intel-gfx] [PATCH] drm/i915: Revert shrinker changes from "Track unbound pages"
Date: Thu, 10 Jan 2013 18:03:38 +0100 [thread overview]
Message-ID: <20130110170338.GE5737@phenom.ffwll.local> (raw)
In-Reply-To: <6c3329$820h5e@orsmga002.jf.intel.com>
On Thu, Jan 10, 2013 at 04:58:30PM +0000, Chris Wilson wrote:
> On Thu, 10 Jan 2013 18:03:00 +0100, Daniel Vetter <daniel.vetter@ffwll.ch> wrote:
> > This partially reverts
> >
> > commit 6c085a728cf000ac1865d66f8c9b52935558b328
> > Author: Chris Wilson <chris@chris-wilson.co.uk>
> > Date: Mon Aug 20 11:40:46 2012 +0200
> >
> > drm/i915: Track unbound pages
> >
> > Closer inspection of that patch revealed a bunch of unrelated changes
> > in the shrinker:
> > - The shrinker count is now in pages instead of objects.
> > - For counting the shrinkable objects the old code only looked at the
> > inactive list, the new code looks at all bounds objects (including
> > pinned ones). That is obviously in addition to the new unbound list.
> > - The shrinker cound is no longer scaled with
> > sysctl_vfs_cache_pressure. Note though that with the default tuning
> > value of vfs_cache_pressue = 100 this doesn't affect the shrinker
> > behaviour.
> > - When actually shrinking objects, the old code first dropped
> > purgeable objects, then normal (inactive) objects. Only then did it,
> > in a last-ditch effort idle the gpu and evict everything. The new
> > code omits the intermediate step of evicting normal inactive
> > objects.
> >
> > Safe for the first change, which seems benign, and the shrinker count
> > scaling, which is a bit a different story, the endresult of all these
> > changes is that the shrinker is _much_ more likely to fall back to the
> > last-ditch resort of idling the gpu and evicting everything. The old
> > code could only do that if something else evicted lots of objects
> > meanwhile (since without any other changes the nr_to_scan will be
> > smaller than the object count).
> >
> > Reverting the vfs_cache_pressure behaviour itself is a bit bogus: Only
> > dentry/inode object caches should scale their shrinker counts with
> > vfs_cache_pressure. Originally I've had that change reverted, too. But
> > Chris Wilson insisted that it's too bogus and shouldn't again see the
> > light of day.
> >
> > Hence revert all these other changes and restore the old shrinker
> > behaviour, with the minor adjustment that we now first scan the
> > unbound list, then the inactive list for each object category
> > (purgeable or normal).
> >
> > A similar patch has been tested by a few people affected by the gen4/5
> > hangs which started to appear in 3.7, which some people bisected to
> > the "drm/i915: Track unbound pages" commit. But just disabling the
> > unbound logic alone didn't change things at all.
> >
> > Note that this patch doesn't fix the referenced bugs, it only hides
> > the underlying bug(s) well enough to restore pre-3.7 behaviour. The
> > key to achieve that is to massively reduce the likelyhood of going
> > into a full gpu stall and evicting everything.
> >
> > v2: Reword commit message a bit, taking Chris Wilson's comment into
> > account.
> >
> > v3: On Chris Wilson's insistency, do not reinstate the rather bogus
> > vfs_cache_pressure change.
> >
> > Tested-by: Greg KH <gregkh@linuxfoundation.org>
> > Tested-by: Dave Kleikamp <dave.kleikamp@oracle.com>
> > References: https://bugs.freedesktop.org/show_bug.cgi?id=55984
> > References: https://bugs.freedesktop.org/show_bug.cgi?id=57122
> > References: https://bugs.freedesktop.org/show_bug.cgi?id=56916
> > References: https://bugs.freedesktop.org/show_bug.cgi?id=57136
> > Cc: Chris Wilson <chris@chris-wilson.co.uk>
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Daniel Vetter <daniel.vetter@ffwll.ch>
>
> Acked-by: Chris Wilson <chris@chris-wilson.co.uk>
Picked up for -fixes.
> I just hope the clue bat descends soonest before we find another way of
> triggering the spurious hangs.
Indeed.
-Daniel
--
Daniel Vetter
Software Engineer, Intel Corporation
+41 (0) 79 365 57 48 - http://blog.ffwll.ch
prev parent reply other threads:[~2013-01-10 17:03 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2013-01-09 17:57 [PATCH] drm/i915: Revert shrinker changes from "Track unbound pages" Daniel Vetter
2013-01-10 0:23 ` Chris Wilson
2013-01-10 9:34 ` Daniel Vetter
2013-01-10 10:21 ` Daniel Vetter
2013-01-10 11:37 ` Daniel Vetter
2013-01-10 17:03 ` Daniel Vetter
2013-01-10 16:58 ` Chris Wilson
2013-01-10 17:03 ` Daniel Vetter [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20130110170338.GE5737@phenom.ffwll.local \
--to=daniel@ffwll.ch \
--cc=chris@chris-wilson.co.uk \
--cc=daniel.vetter@ffwll.ch \
--cc=intel-gfx@lists.freedesktop.org \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox