From: Nick Piggin <npiggin@kernel.dk>
To: Dave Chinner <david@fromorbit.com>
Cc: Nick Piggin <npiggin@kernel.dk>,
Christoph Hellwig <hch@infradead.org>,
linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH 17/18] fs: icache remove inode_lock
Date: Thu, 14 Oct 2010 20:06:09 +1100 [thread overview]
Message-ID: <20101014090609.GB3144@amd> (raw)
In-Reply-To: <20101013232319.GZ4681@dastard>
On Thu, Oct 14, 2010 at 10:23:19AM +1100, Dave Chinner wrote:
> On Wed, Oct 13, 2010 at 11:30:08PM +1100, Nick Piggin wrote:
> > On Wed, Oct 13, 2010 at 07:25:52AM -0400, Christoph Hellwig wrote:
> > > On Wed, Oct 13, 2010 at 06:20:58PM +1100, Nick Piggin wrote:
> > > > Unfortunate timing that everybody is suddenly interested in the
> > > > scalability work :)
> > >
> > > People have been interested for a long time. It's just that we finally
> > > made forward progress to get parts of it shape for merging, which should
> > > have been done long time ago.
> >
> > It has been pretty close, IMO, it just needed some more reviewers, which
> > it only just got really.
>
> Going back a couple of weeks, it seemed as far away from inclusion
> as when you first posted the series - there had been no substanital
> review and nobody wanting to review it in it's current form. Then
> I'd heard you were travelling indefinitely.....
Well it was a few weeks. Not great timing, but as I said, I didn't
want to dump the latest patches and then disappear.
There were actually several (google and intel) people testing things.
Last I heard from you (from the first constructive review I had), you
were skeptical about performance and scalability required and thought
it would be a better idea to reduce lock widths and things incrementally
(which I believe will lead to much more overall churn and confusion for
maintaining between different kernel versions).
Anyway, no real harm done and I thank you for picking things up, so
I'm still working through what you've done atm.
>
> There is stuff in the vfs-scale tree that is somewhat controversial
> and had not been discussed satisfactorily - the lock ordering
> (resulting in trylocks everywhere), the shrinker API change, the
> writeback LRU changes, the zone reclaim changes, etc - and some of
> them even have alternative proposals for fixing the algorithmic
> deficiencies. Nobody was going to review or accept that as one big
> lump.
I think what is absolutely needed is a final(ish) "lump", so we can
actually see what we're working towards and whether it is the right
approach.
Shrinker and zone reclaim is definitely needed. It is needed for NUMA
scalability and locality of reclaim, and also for container and directed
dentry/inode reclaim. Google have a very similar patch and they've said
this is needed (and I already know it is needed for scalability on
large NUMA -- SGI were complaining about this nearly 5 years ago IIRC).
So that is _definitely_ going to be needed.
Store-free path walking is definitely needed, so we need to do RCU inodes.
With RCU inodes, the optimal locking protocols change quite a bit --
Trylocks are a side effect of my conservative approach to establishing
a lock order, and making incremental changes in locking protocol which
are supposed to be easy to follow and verify. Then comes optimisation
of the locking after the "correctness" part of it. It can be taken even
further and *most* of the icache trylocks removed using RCU on a few
more of the lists.
Now I don't want to necessarily merge it all at once, but I do need an
overview of it, most certainly.
A few trylocks confined to core inode handling (and the other inode
locks are not exposed to filesystems at all -- unlike inode_lock today)
is not what I'd call a maintainence nightmare. In fact, IMO it is
cleaner locking especially for filesystems than today.
With a structure like icache, where icache can be accessed down, via
the cache management structures or up, via references to inodes,
trylocks are not unusual unless RCU is extensively used. As long as
I'm careful, I don't see a problem at all -- XFS similarly has some
trylocks.
> There's been review now because I went and did what the potential
> reviewers were asking for - break the series into smaller, more
> easily reviewable and verifiable chunks. As a result, I think we're
> close to the end of the review cycle for the inode_lock breakup now.
> I think the code is now much cleaner and more maintainable than what I
> originally pulled from the vfs-scale tree, and it still provides the
> same gains and ability to be converted to RCU algorithms in the
> future.
>
> Hence, IMO, the current vfs-scale tree needs to be treated as a
> prototype, not as a finished product. It demonstrates the the path
> we need to follow to move forward, as well as the gains we will get
> as we move in that direction, but the code in that tree is not
> guaranteed a free pass into the mainline tree.
It is really pretty close, and while *you* have some disagreements,
it has had some reviews from other people (including Linus) who actually
agree with most of it and agree that scalability is needed.
I am fine with continuing to take suggestions that I agree with, and
integrating them into the tree and begin to start pushing chunks out
the bottom of the stack, but I would like to keep things together in
my upstream tree.
> > etc. I had to really learn most of the code from scratch to get this far
> > and got quite little constructive review really until recently.
>
> Which seems to be another good reason for treating the tree as prototype
> code.
It's much past a prototype. While the patches need some more cleanup
and review still, the final end result gives a tree with almost no
global cachelines in the entire vfs, including path walking. Things
like path walks are nearly 50% faster single threaded, and perfectly
scalable. Linus actually wants the store-free path walk stuff
_before_ any of the other things, if that gives you an idea of where
other people are putting the priority of the patches.
Thanks,
Nick
next prev parent reply other threads:[~2010-10-14 9:06 UTC|newest]
Thread overview: 162+ messages / expand[flat|nested] mbox.gz Atom feed top
2010-10-08 5:21 fs: Inode cache scalability V2 Dave Chinner
2010-10-08 5:21 ` [PATCH 01/18] kernel: add bl_list Dave Chinner
2010-10-08 8:18 ` Andi Kleen
2010-10-08 10:33 ` Dave Chinner
2010-10-08 5:21 ` [PATCH 02/18] fs: Convert nr_inodes and nr_unused to per-cpu counters Dave Chinner
2010-10-08 7:01 ` Christoph Hellwig
2010-10-08 5:21 ` [PATCH 03/18] fs: keep inode with backing-dev Dave Chinner
2010-10-08 7:01 ` Christoph Hellwig
2010-10-08 7:27 ` Dave Chinner
2010-10-08 5:21 ` [PATCH 04/18] fs: Implement lazy LRU updates for inodes Dave Chinner
2010-10-08 7:08 ` Christoph Hellwig
2010-10-08 7:31 ` Dave Chinner
2010-10-08 9:08 ` Al Viro
2010-10-08 9:51 ` Dave Chinner
2010-10-08 5:21 ` [PATCH 05/18] fs: inode split IO and LRU lists Dave Chinner
2010-10-08 7:14 ` Christoph Hellwig
2010-10-08 7:38 ` Dave Chinner
2010-10-08 9:16 ` Al Viro
2010-10-08 9:58 ` Dave Chinner
2010-10-08 5:21 ` [PATCH 06/18] fs: Clean up inode reference counting Dave Chinner
2010-10-08 7:20 ` Christoph Hellwig
2010-10-08 7:46 ` Dave Chinner
2010-10-08 8:15 ` Christoph Hellwig
2010-10-08 5:21 ` [PATCH 07/18] exofs: use iput() for inode reference count decrements Dave Chinner
2010-10-08 7:21 ` Christoph Hellwig
2010-10-16 7:56 ` Nick Piggin
2010-10-16 16:29 ` Christoph Hellwig
2010-10-17 15:41 ` Boaz Harrosh
2010-10-08 5:21 ` [PATCH 08/18] fs: add inode reference coutn read accessor Dave Chinner
2010-10-08 7:24 ` Christoph Hellwig
2010-10-08 5:21 ` [PATCH 09/18] fs: rework icount to be a locked variable Dave Chinner
2010-10-08 7:27 ` Christoph Hellwig
2010-10-08 7:50 ` Dave Chinner
2010-10-08 8:17 ` Christoph Hellwig
2010-10-08 13:16 ` Chris Mason
2010-10-08 9:32 ` Al Viro
2010-10-08 10:15 ` Dave Chinner
2010-10-08 13:14 ` Chris Mason
2010-10-08 13:53 ` Christoph Hellwig
2010-10-08 14:09 ` Dave Chinner
2010-10-08 5:21 ` [PATCH 10/18] fs: Factor inode hash operations into functions Dave Chinner
2010-10-08 7:29 ` Christoph Hellwig
2010-10-08 9:41 ` Al Viro
2010-10-08 5:21 ` [PATCH 11/18] fs: Introduce per-bucket inode hash locks Dave Chinner
2010-10-08 7:33 ` Christoph Hellwig
2010-10-08 7:51 ` Dave Chinner
2010-10-08 9:49 ` Al Viro
2010-10-08 9:51 ` Christoph Hellwig
2010-10-08 13:43 ` Christoph Hellwig
2010-10-08 14:17 ` Dave Chinner
2010-10-08 18:54 ` Christoph Hellwig
2010-10-16 7:57 ` Nick Piggin
2010-10-16 16:16 ` Christoph Hellwig
2010-10-16 17:12 ` Nick Piggin
2010-10-17 0:45 ` Christoph Hellwig
2010-10-17 2:06 ` Nick Piggin
2010-10-17 0:46 ` Dave Chinner
2010-10-17 2:25 ` Nick Piggin
2010-10-18 16:16 ` Andi Kleen
2010-10-18 16:21 ` Christoph Hellwig
2010-10-19 7:00 ` Nick Piggin
2010-10-19 16:50 ` Christoph Hellwig
2010-10-20 3:11 ` Nick Piggin
2010-10-24 15:44 ` Thomas Gleixner
2010-10-24 21:17 ` Nick Piggin
2010-10-25 4:41 ` Thomas Gleixner
2010-10-25 7:04 ` Thomas Gleixner
2010-10-26 0:12 ` Nick Piggin
2010-10-26 0:06 ` Nick Piggin
2010-10-08 5:21 ` [PATCH 12/18] fs: add a per-superblock lock for the inode list Dave Chinner
2010-10-08 7:35 ` Christoph Hellwig
2010-10-08 5:21 ` [PATCH 13/18] fs: split locking of inode writeback and LRU lists Dave Chinner
2010-10-08 7:42 ` Christoph Hellwig
2010-10-08 8:00 ` Dave Chinner
2010-10-08 8:18 ` Christoph Hellwig
2010-10-16 7:57 ` Nick Piggin
2010-10-16 16:20 ` Christoph Hellwig
2010-10-16 17:19 ` Nick Piggin
2010-10-17 1:00 ` Dave Chinner
2010-10-17 2:20 ` Nick Piggin
2010-10-08 5:21 ` [PATCH 14/18] fs: Protect inode->i_state with th einode->i_lock Dave Chinner
2010-10-08 7:49 ` Christoph Hellwig
2010-10-08 8:04 ` Dave Chinner
2010-10-08 8:18 ` Christoph Hellwig
2010-10-16 7:57 ` Nick Piggin
2010-10-16 16:19 ` Christoph Hellwig
2010-10-09 8:05 ` Christoph Hellwig
2010-10-09 14:52 ` Matthew Wilcox
2010-10-10 2:01 ` Dave Chinner
2010-10-08 5:21 ` [PATCH 15/18] fs: introduce a per-cpu last_ino allocator Dave Chinner
2010-10-08 7:53 ` Christoph Hellwig
2010-10-08 8:05 ` Dave Chinner
2010-10-08 8:22 ` Andi Kleen
2010-10-08 8:44 ` Christoph Hellwig
2010-10-08 9:58 ` Al Viro
2010-10-08 10:09 ` Andi Kleen
2010-10-08 10:19 ` Al Viro
2010-10-08 10:20 ` Eric Dumazet
2010-10-08 9:56 ` Al Viro
2010-10-08 10:03 ` Christoph Hellwig
2010-10-08 10:20 ` Eric Dumazet
2010-10-08 13:48 ` Christoph Hellwig
2010-10-08 14:06 ` Eric Dumazet
2010-10-08 19:10 ` Christoph Hellwig
2010-10-09 17:14 ` Matthew Wilcox
2010-10-16 7:57 ` Nick Piggin
2010-10-16 16:22 ` Christoph Hellwig
2010-10-16 17:21 ` Nick Piggin
2010-10-08 5:21 ` [PATCH 16/18] fs: Make iunique independent of inode_lock Dave Chinner
2010-10-08 7:55 ` Christoph Hellwig
2010-10-08 8:06 ` Dave Chinner
2010-10-08 8:19 ` Christoph Hellwig
2010-10-08 5:21 ` [PATCH 17/18] fs: icache remove inode_lock Dave Chinner
2010-10-08 8:03 ` Christoph Hellwig
2010-10-08 8:09 ` Dave Chinner
2010-10-13 7:20 ` Nick Piggin
2010-10-13 7:27 ` Nick Piggin
2010-10-13 11:28 ` Christoph Hellwig
2010-10-13 12:03 ` Nick Piggin
2010-10-13 12:20 ` Christoph Hellwig
2010-10-13 12:25 ` Nick Piggin
2010-10-13 10:42 ` Eric Dumazet
2010-10-13 12:07 ` Nick Piggin
2010-10-13 11:25 ` Christoph Hellwig
2010-10-13 12:30 ` Nick Piggin
2010-10-13 23:23 ` Dave Chinner
2010-10-14 9:06 ` Nick Piggin [this message]
2010-10-14 9:13 ` Nick Piggin
2010-10-14 14:41 ` Christoph Hellwig
2010-10-15 0:14 ` Nick Piggin
2010-10-15 3:13 ` Dave Chinner
2010-10-15 3:30 ` Nick Piggin
2010-10-15 3:44 ` Nick Piggin
2010-10-15 6:41 ` Nick Piggin
2010-10-15 10:59 ` Dave Chinner
2010-10-15 13:03 ` Nick Piggin
2010-10-15 13:29 ` Nick Piggin
2010-10-15 17:33 ` Nick Piggin
2010-10-15 17:52 ` Christoph Hellwig
2010-10-15 18:02 ` Nick Piggin
2010-10-15 18:14 ` Nick Piggin
2010-10-16 2:09 ` Nick Piggin
2010-10-15 14:11 ` Nick Piggin
2010-10-15 20:50 ` Nick Piggin
2010-10-15 20:56 ` Nick Piggin
2010-10-15 4:04 ` Nick Piggin
2010-10-15 11:33 ` Dave Chinner
2010-10-15 13:14 ` Nick Piggin
2010-10-15 15:38 ` Nick Piggin
2010-10-16 7:57 ` Nick Piggin
2010-10-08 5:21 ` [PATCH 18/18] fs: Reduce inode I_FREEING and factor inode disposal Dave Chinner
2010-10-08 8:11 ` Christoph Hellwig
2010-10-08 10:18 ` Al Viro
2010-10-08 10:52 ` Dave Chinner
2010-10-08 12:10 ` Al Viro
2010-10-08 13:55 ` Dave Chinner
2010-10-09 17:22 ` Matthew Wilcox
2010-10-09 8:08 ` [PATCH 19/18] fs: split __inode_add_to_list Christoph Hellwig
2010-10-12 10:47 ` Dave Chinner
2010-10-12 11:31 ` Christoph Hellwig
2010-10-12 12:05 ` Dave Chinner
2010-10-09 11:18 ` [PATCH 20/18] fs: do not assign default i_ino in new_inode Christoph Hellwig
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20101014090609.GB3144@amd \
--to=npiggin@kernel.dk \
--cc=david@fromorbit.com \
--cc=hch@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).