From: Trond Myklebust <trondmy@hammerspace.com>
To: "bcodding@redhat.com" <bcodding@redhat.com>
Cc: "linux-nfs@vger.kernel.org" <linux-nfs@vger.kernel.org>,
"neilb@suse.de" <neilb@suse.de>
Subject: Re: [PATCH v7 19/21] NFS: Convert readdir page cache to use a cookie based index
Date: Fri, 25 Feb 2022 13:11:41 +0000 [thread overview]
Message-ID: <7252c8e5c9e2a2729758ad64e0e6e75cd5e0d730.camel@hammerspace.com> (raw)
In-Reply-To: <AA3FF18A-38CE-4F5A-87FE-6C235C6CD9BD@redhat.com>
On Fri, 2022-02-25 at 07:33 -0500, Benjamin Coddington wrote:
> On 24 Feb 2022, at 23:25, Trond Myklebust wrote:
>
> > On Fri, 2022-02-25 at 14:17 +1100, NeilBrown wrote:
>
> > > I haven't looked at the code recently so this might not be 100%
> > > accurate,
> > > but XArray generally assumes that pages are often adjacent. They
> > > don't
> > > have to be, but there is a cost. It uses a multi-level array
> > > with 9 bits
> > > per level. At each level there are a whole number of pages for
> > > indexes
> > > to the next level.
> > >
> > > If there are two entries, that are 2^45 separate, that is 5
> > > levels of
> > > indexing that cannot be shared. So the path to one entry is 5
> > > pages,
> > > each of which contains a single pointer. The path to the other
> > > entry is
> > > a separate set of 5 pages.
> > >
> > > So worst case, the index would be about 64/9 or 7 times the size
> > > of the
> > > data. As the number of data pages increases, this would shrink
> > > slightly,
> > > but I suspect you wouldn't get below a factor of 3 before you
> > > fill up all
> > > of your memory.
>
> Yikes!
>
> > If the problem is just the range, then that is trivial to fix: we
> > can
> > just use xxhash32(), and take the hit of more collisions. However
> > if
> > the problem is the access pattern, then I have serious questions
> > about
> > the choice of implementation for the page cache. If the cache can't
> > support file random access, then we're barking up the wrong tree on
> > the
> > wrong continent.
>
> I'm guessing the issue might be "get next", which for an "array" is
> probably
> the operation tested for "perform well". We're not doing any of
> that, we're
> directly addressing pages with our hashed index.
>
> > Either way, I see avoiding linear searches for cookies as a benefit
> > that is worth pursuing.
>
> Me too. What about just kicking the seekdir users up into the second
> half
> of the index, to use xxhash32() up there. Everyone else can hang out
> in the
> bottom half with dense indexes and help each other out.
>
> The vast majority of readdir() use is going to be short listings
> traversed
> in order. The memory inflation created by a process that needs to
> walk a
> tree and for every two pages of readdir data require 10 pages of
> indexes
> seems pretty extreme.
>
> Ben
>
#define NFS_READDIR_COOKIE_MASK (U32_MAX >> 14)
/*
* Hash algorithm allowing content addressible access to sequences
* of directory cookies. Content is addressed by the value of the
* cookie index of the first readdir entry in a page.
*
* The xxhash algorithm is chosen because it is fast, and is supposed
* to result in a decent flat distribution of hashes.
*
* We then select only the first 18 bits to avoid issues with excessive
* memory use for the page cache XArray. 18 bits should allow the caching
* of 262144 pages of sequences of readdir entries. Since each page holds
* 127 readdir entries for a typical 64-bit system, that works out to a
* cache of ~ 33 million entries per directory.
*/
static pgoff_t nfs_readdir_page_cookie_hash(u64 cookie)
{
if (cookie == 0)
return 0;
return xxhash(&cookie, sizeof(cookie), 0) & NFS_READDIR_COOKIE_MASK;
}
So no, this is not a show-stopper.
--
Trond Myklebust
Linux NFS client maintainer, Hammerspace
trond.myklebust@hammerspace.com
next prev parent reply other threads:[~2022-02-25 13:11 UTC|newest]
Thread overview: 57+ messages / expand[flat|nested] mbox.gz Atom feed top
2022-02-23 21:12 [PATCH v7 00/21] Readdir improvements trondmy
2022-02-23 21:12 ` [PATCH v7 01/21] NFS: constify nfs_server_capable() and nfs_have_writebacks() trondmy
2022-02-23 21:12 ` [PATCH v7 02/21] NFS: Trace lookup revalidation failure trondmy
2022-02-23 21:12 ` [PATCH v7 03/21] NFS: Use kzalloc() to avoid initialising the nfs_open_dir_context trondmy
2022-02-23 21:12 ` [PATCH v7 04/21] NFS: Calculate page offsets algorithmically trondmy
2022-02-23 21:12 ` [PATCH v7 05/21] NFS: Store the change attribute in the directory page cache trondmy
2022-02-23 21:12 ` [PATCH v7 06/21] NFS: If the cookie verifier changes, we must invalidate the " trondmy
2022-02-23 21:12 ` [PATCH v7 07/21] NFS: Don't re-read the entire page cache to find the next cookie trondmy
2022-02-23 21:12 ` [PATCH v7 08/21] NFS: Adjust the amount of readahead performed by NFS readdir trondmy
2022-02-23 21:12 ` [PATCH v7 09/21] NFS: Simplify nfs_readdir_xdr_to_array() trondmy
2022-02-23 21:12 ` [PATCH v7 10/21] NFS: Reduce use of uncached readdir trondmy
2022-02-23 21:12 ` [PATCH v7 11/21] NFS: Improve heuristic for readdirplus trondmy
2022-02-23 21:12 ` [PATCH v7 12/21] NFS: Don't ask for readdirplus unless it can help nfs_getattr() trondmy
2022-02-23 21:12 ` [PATCH v7 13/21] NFSv4: Ask for a full XDR buffer of readdir goodness trondmy
2022-02-23 21:12 ` [PATCH v7 14/21] NFS: Readdirplus can't help lookup for case insensitive filesystems trondmy
2022-02-23 21:12 ` [PATCH v7 15/21] NFS: Don't request readdirplus when revalidation was forced trondmy
2022-02-23 21:13 ` [PATCH v7 16/21] NFS: Add basic readdir tracing trondmy
2022-02-23 21:13 ` [PATCH v7 17/21] NFS: Trace effects of readdirplus on the dcache trondmy
2022-02-23 21:13 ` [PATCH v7 18/21] NFS: Trace effects of the readdirplus heuristic trondmy
2022-02-23 21:13 ` [PATCH v7 19/21] NFS: Convert readdir page cache to use a cookie based index trondmy
2022-02-23 21:13 ` [PATCH v7 20/21] NFS: Fix up forced readdirplus trondmy
2022-02-23 21:13 ` [PATCH v7 21/21] NFS: Remove unnecessary cache invalidations for directories trondmy
2022-02-24 17:31 ` [PATCH v7 19/21] NFS: Convert readdir page cache to use a cookie based index Benjamin Coddington
2022-02-25 2:33 ` Trond Myklebust
2022-02-25 3:17 ` NeilBrown
2022-02-25 4:25 ` Trond Myklebust
2022-02-25 12:33 ` Benjamin Coddington
2022-02-25 13:11 ` Trond Myklebust [this message]
2022-02-24 15:53 ` [PATCH v7 16/21] NFS: Add basic readdir tracing Benjamin Coddington
2022-02-25 2:35 ` Trond Myklebust
2022-02-24 16:55 ` [PATCH v7 10/21] NFS: Reduce use of uncached readdir Anna Schumaker
2022-02-25 4:07 ` Trond Myklebust
2022-02-24 16:30 ` [PATCH v7 08/21] NFS: Adjust the amount of readahead performed by NFS readdir Anna Schumaker
2022-02-24 16:18 ` [PATCH v7 06/21] NFS: If the cookie verifier changes, we must invalidate the page cache Anna Schumaker
2022-02-24 14:53 ` [PATCH v7 05/21] NFS: Store the change attribute in the directory " Benjamin Coddington
2022-02-25 2:26 ` Trond Myklebust
2022-02-25 3:51 ` Trond Myklebust
2022-02-25 11:38 ` Benjamin Coddington
2022-02-25 13:10 ` Trond Myklebust
2022-02-25 13:26 ` Trond Myklebust
2022-02-25 14:44 ` Benjamin Coddington
2022-02-25 15:18 ` Trond Myklebust
2022-02-25 15:34 ` Benjamin Coddington
2022-02-25 20:23 ` Benjamin Coddington
2022-02-25 20:28 ` Benjamin Coddington
2022-02-25 20:41 ` Trond Myklebust
2022-02-25 22:04 ` Benjamin Coddington
2022-02-25 22:29 ` Trond Myklebust
2022-02-24 14:15 ` [PATCH v7 04/21] NFS: Calculate page offsets algorithmically Benjamin Coddington
2022-02-25 2:11 ` Trond Myklebust
2022-02-25 11:28 ` Benjamin Coddington
2022-02-25 12:44 ` Trond Myklebust
2022-02-24 14:14 ` [PATCH v7 02/21] NFS: Trace lookup revalidation failure Benjamin Coddington
2022-02-25 2:09 ` Trond Myklebust
2022-02-24 12:25 ` [PATCH v7 00/21] Readdir improvements David Wysochanski
2022-02-25 4:00 ` Trond Myklebust
2022-02-24 15:07 ` David Wysochanski
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7252c8e5c9e2a2729758ad64e0e6e75cd5e0d730.camel@hammerspace.com \
--to=trondmy@hammerspace.com \
--cc=bcodding@redhat.com \
--cc=linux-nfs@vger.kernel.org \
--cc=neilb@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox