From: Trond Myklebust <trondmy@hammerspace.com>
To: "neilb@suse.de" <neilb@suse.de>
Cc: "linux-nfs@vger.kernel.org" <linux-nfs@vger.kernel.org>,
"bcodding@redhat.com" <bcodding@redhat.com>
Subject: Re: [PATCH v7 19/21] NFS: Convert readdir page cache to use a cookie based index
Date: Fri, 25 Feb 2022 04:25:03 +0000 [thread overview]
Message-ID: <4f1a9a7b5e3ac59e365c5e40ee146ceb0f4e1429.camel@hammerspace.com> (raw)
In-Reply-To: <164575906990.4638.4113048743095971193@noble.neil.brown.name>
On Fri, 2022-02-25 at 14:17 +1100, NeilBrown wrote:
> On Fri, 25 Feb 2022, Trond Myklebust wrote:
> > On Thu, 2022-02-24 at 12:31 -0500, Benjamin Coddington wrote:
> > >
> > > "The XArray implementation is efficient when the indices used are
> > > densely
> > > clustered; hashing the object and using the hash as the index
> > > will
> > > not
> > > perform well."
> > >
> > > However, the "not perform well" may be orders of magnitude
> > > smaller
> > > than
> > > anthing like RPC. Do you have concerns about this?
> >
> > What is the difference between this workload and a random access
> > database workload?
>
> Probably the range of expected addresses.
> If I understand the proposal correctly, the page addresses in this
> workload could be any 64bit number.
> For a large database, it would be at most 52 bits (assuming 64bits
> worth
> of bytes), and very likely substantially smaller - maybe 40 bits for
> a
> really really big database.
>
> >
> > If the XArray is incapable of dealing with random access, then we
> > should never have chosen it for the page cache. I'm therefore
> > assuming
> > that either the above comment is referring to micro-optimisations
> > that
> > don't matter much with these workloads, or else that the plan is to
> > replace the XArray with something more appropriate for a page cache
> > workload.
>
> I haven't looked at the code recently so this might not be 100%
> accurate, but XArray generally assumes that pages are often adjacent.
> They don't have to be, but there is a cost.
> It uses a multi-level array with 9 bits per level. At each level
> there
> are a whole number of pages for indexes to the next level.
>
> If there are two entries, that are 2^45 separate, that is 5 levels of
> indexing that cannot be shared. So the path to one entry is 5 pages,
> each of which contains a single pointer. The path to the other entry
> is
> a separate set of 5 pages.
>
> So worst case, the index would be about 64/9 or 7 times the size of
> the
> data. As the number of data pages increases, this would shrink
> slightly, but I suspect you wouldn't get below a factor of 3 before
> you
> fill up all of your memory.
>
If the problem is just the range, then that is trivial to fix: we can
just use xxhash32(), and take the hit of more collisions. However if
the problem is the access pattern, then I have serious questions about
the choice of implementation for the page cache. If the cache can't
support file random access, then we're barking up the wrong tree on the
wrong continent.
Either way, I see avoiding linear searches for cookies as a benefit
that is worth pursuing.
--
Trond Myklebust
Linux NFS client maintainer, Hammerspace
trond.myklebust@hammerspace.com
next prev parent reply other threads:[~2022-02-25 4:25 UTC|newest]
Thread overview: 57+ messages / expand[flat|nested] mbox.gz Atom feed top
2022-02-23 21:12 [PATCH v7 00/21] Readdir improvements trondmy
2022-02-23 21:12 ` [PATCH v7 01/21] NFS: constify nfs_server_capable() and nfs_have_writebacks() trondmy
2022-02-23 21:12 ` [PATCH v7 02/21] NFS: Trace lookup revalidation failure trondmy
2022-02-23 21:12 ` [PATCH v7 03/21] NFS: Use kzalloc() to avoid initialising the nfs_open_dir_context trondmy
2022-02-23 21:12 ` [PATCH v7 04/21] NFS: Calculate page offsets algorithmically trondmy
2022-02-23 21:12 ` [PATCH v7 05/21] NFS: Store the change attribute in the directory page cache trondmy
2022-02-23 21:12 ` [PATCH v7 06/21] NFS: If the cookie verifier changes, we must invalidate the " trondmy
2022-02-23 21:12 ` [PATCH v7 07/21] NFS: Don't re-read the entire page cache to find the next cookie trondmy
2022-02-23 21:12 ` [PATCH v7 08/21] NFS: Adjust the amount of readahead performed by NFS readdir trondmy
2022-02-23 21:12 ` [PATCH v7 09/21] NFS: Simplify nfs_readdir_xdr_to_array() trondmy
2022-02-23 21:12 ` [PATCH v7 10/21] NFS: Reduce use of uncached readdir trondmy
2022-02-23 21:12 ` [PATCH v7 11/21] NFS: Improve heuristic for readdirplus trondmy
2022-02-23 21:12 ` [PATCH v7 12/21] NFS: Don't ask for readdirplus unless it can help nfs_getattr() trondmy
2022-02-23 21:12 ` [PATCH v7 13/21] NFSv4: Ask for a full XDR buffer of readdir goodness trondmy
2022-02-23 21:12 ` [PATCH v7 14/21] NFS: Readdirplus can't help lookup for case insensitive filesystems trondmy
2022-02-23 21:12 ` [PATCH v7 15/21] NFS: Don't request readdirplus when revalidation was forced trondmy
2022-02-23 21:13 ` [PATCH v7 16/21] NFS: Add basic readdir tracing trondmy
2022-02-23 21:13 ` [PATCH v7 17/21] NFS: Trace effects of readdirplus on the dcache trondmy
2022-02-23 21:13 ` [PATCH v7 18/21] NFS: Trace effects of the readdirplus heuristic trondmy
2022-02-23 21:13 ` [PATCH v7 19/21] NFS: Convert readdir page cache to use a cookie based index trondmy
2022-02-23 21:13 ` [PATCH v7 20/21] NFS: Fix up forced readdirplus trondmy
2022-02-23 21:13 ` [PATCH v7 21/21] NFS: Remove unnecessary cache invalidations for directories trondmy
2022-02-24 17:31 ` [PATCH v7 19/21] NFS: Convert readdir page cache to use a cookie based index Benjamin Coddington
2022-02-25 2:33 ` Trond Myklebust
2022-02-25 3:17 ` NeilBrown
2022-02-25 4:25 ` Trond Myklebust [this message]
2022-02-25 12:33 ` Benjamin Coddington
2022-02-25 13:11 ` Trond Myklebust
2022-02-24 15:53 ` [PATCH v7 16/21] NFS: Add basic readdir tracing Benjamin Coddington
2022-02-25 2:35 ` Trond Myklebust
2022-02-24 16:55 ` [PATCH v7 10/21] NFS: Reduce use of uncached readdir Anna Schumaker
2022-02-25 4:07 ` Trond Myklebust
2022-02-24 16:30 ` [PATCH v7 08/21] NFS: Adjust the amount of readahead performed by NFS readdir Anna Schumaker
2022-02-24 16:18 ` [PATCH v7 06/21] NFS: If the cookie verifier changes, we must invalidate the page cache Anna Schumaker
2022-02-24 14:53 ` [PATCH v7 05/21] NFS: Store the change attribute in the directory " Benjamin Coddington
2022-02-25 2:26 ` Trond Myklebust
2022-02-25 3:51 ` Trond Myklebust
2022-02-25 11:38 ` Benjamin Coddington
2022-02-25 13:10 ` Trond Myklebust
2022-02-25 13:26 ` Trond Myklebust
2022-02-25 14:44 ` Benjamin Coddington
2022-02-25 15:18 ` Trond Myklebust
2022-02-25 15:34 ` Benjamin Coddington
2022-02-25 20:23 ` Benjamin Coddington
2022-02-25 20:28 ` Benjamin Coddington
2022-02-25 20:41 ` Trond Myklebust
2022-02-25 22:04 ` Benjamin Coddington
2022-02-25 22:29 ` Trond Myklebust
2022-02-24 14:15 ` [PATCH v7 04/21] NFS: Calculate page offsets algorithmically Benjamin Coddington
2022-02-25 2:11 ` Trond Myklebust
2022-02-25 11:28 ` Benjamin Coddington
2022-02-25 12:44 ` Trond Myklebust
2022-02-24 14:14 ` [PATCH v7 02/21] NFS: Trace lookup revalidation failure Benjamin Coddington
2022-02-25 2:09 ` Trond Myklebust
2022-02-24 12:25 ` [PATCH v7 00/21] Readdir improvements David Wysochanski
2022-02-25 4:00 ` Trond Myklebust
2022-02-24 15:07 ` David Wysochanski
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4f1a9a7b5e3ac59e365c5e40ee146ceb0f4e1429.camel@hammerspace.com \
--to=trondmy@hammerspace.com \
--cc=bcodding@redhat.com \
--cc=linux-nfs@vger.kernel.org \
--cc=neilb@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox