From: Dave Chinner <david@fromorbit.com>
To: NeilBrown <neilb@suse.de>
Cc: Chuck Lever III <chuck.lever@oracle.com>,
Trond Myklebust <trondmy@hammerspace.com>,
Linux NFS Mailing List <linux-nfs@vger.kernel.org>,
"jlayton@kernel.org" <jlayton@kernel.org>
Subject: Re: knfsd performance
Date: Thu, 20 Jun 2024 12:29:47 +1000 [thread overview]
Message-ID: <ZnOUG2Nh80vTJXxe@dread.disaster.area> (raw)
In-Reply-To: <171883231568.14261.16495433738354176501@noble.neil.brown.name>
On Thu, Jun 20, 2024 at 07:25:15AM +1000, NeilBrown wrote:
> On Wed, 19 Jun 2024, NeilBrown wrote:
> > On Wed, 19 Jun 2024, Dave Chinner wrote:
> > > On Tue, Jun 18, 2024 at 07:54:43PM +0000, Chuck Lever III wrote > On Jun 18, 2024, at 3:50 PM, Trond Myklebust <trondmy@hammerspace.com> wrote:
> > > > >
> > > > > On Tue, 2024-06-18 at 19:39 +0000, Chuck Lever III wrote:
> > > > >>
> > > > >>
> > > > >>> On Jun 18, 2024, at 3:29 PM, Trond Myklebust
> > > > >>> <trondmy@hammerspace.com> wrote:
> > > > >>>
> > > > >>> On Tue, 2024-06-18 at 18:40 +0000, Chuck Lever III wrote:
> > > > >>>>
> > > > >>>>
> > > > >>>>> On Jun 18, 2024, at 2:32 PM, Trond Myklebust
> > > > >>>>> <trondmy@hammerspace.com> wrote:
> > > > >>>>>
> > > > >>>>> I recently back ported Neil's lwq code and sunrpc server
> > > > >>>>> changes to
> > > > >>>>> our
> > > > >>>>> 5.15.130 based kernel in the hope of improving the performance
> > > > >>>>> for
> > > > >>>>> our
> > > > >>>>> data servers.
> > > > >>>>>
> > > > >>>>> Our performance team recently ran a fio workload on a client
> > > > >>>>> that
> > > > >>>>> was
> > > > >>>>> doing 100% NFSv3 reads in O_DIRECT mode over an RDMA connection
> > > > >>>>> (infiniband) against that resulting server. I've attached the
> > > > >>>>> resulting
> > > > >>>>> flame graph from a perf profile run on the server side.
> > > > >>>>>
> > > > >>>>> Is anyone else seeing this massive contention for the spin lock
> > > > >>>>> in
> > > > >>>>> __lwq_dequeue? As you can see, it appears to be dwarfing all
> > > > >>>>> the
> > > > >>>>> other
> > > > >>>>> nfsd activity on the system in question here, being responsible
> > > > >>>>> for
> > > > >>>>> 45%
> > > > >>>>> of all the perf hits.
> > >
> > > Ouch. __lwq_dequeue() runs llist_reverse_order() under a spinlock.
> > >
> > > llist_reverse_order() is an O(n) algorithm involving full length
> > > linked list traversal. IOWs, it's a worst case cache miss algorithm
> > > running under a spin lock. And then consider what happens when
> > > enqueue processing is faster than dequeue processing.
> >
> > My expectation was that if enqueue processing (incoming packets) was
> > faster than dequeue processing (handling NFS requests) then there was a
> > bottleneck elsewhere, and this one wouldn't be relevant.
> >
> > It might be useful to measure how long the queue gets.
>
> Thinking about this some more .... if it did turn out that the queue
> gets long, and maybe even if it didn't, we could reimplement lwq as a
> simple linked list with head and tail pointers.
>
> enqueue would be something like:
>
> new->next = NULL;
> old_tail = xchg(&q->tail, new);
> if (old_tail)
> /* dequeue of old_tail cannot succeed until this assignment completes */
> old_tail->next = new
> else
> q->head = new
>
> dequeue would be
>
> spinlock()
> ret = q->head;
> if (ret) {
> while (ret->next == NULL && cmp_xchg(&q->tail, ret, NULL) != ret)
> /* wait for enqueue of q->tail to complete */
> cpu_relax();
> }
> cmp_xchg(&q->head, ret, ret->next);
> spin_unlock();
That might work, but I suspect that it's still only putting off the
inevitable.
Doing the dequeue purely with atomic operations might be possible,
but it's not immediately obvious to me how to solve both head/tail
race conditions with atomic operations. I can work out an algorithm
that makes enqueue safe against dequeue races (or vice versa), but I
can't also get the logic on the opposite side to also be safe.
I'll let it bounce around my head a bit more...
-Dave.
--
Dave Chinner
david@fromorbit.com
next prev parent reply other threads:[~2024-06-20 2:29 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2024-06-18 18:32 knfsd performance Trond Myklebust
2024-06-18 18:40 ` Chuck Lever III
2024-06-18 19:29 ` Trond Myklebust
2024-06-18 19:39 ` Chuck Lever III
2024-06-18 19:50 ` Trond Myklebust
2024-06-18 19:54 ` Chuck Lever III
2024-06-18 20:16 ` Jeff Layton
2024-06-18 23:17 ` NeilBrown
2024-06-18 23:26 ` Chuck Lever III
2024-06-18 23:33 ` Jeff Layton
2024-06-18 23:51 ` Chuck Lever III
2024-06-19 2:56 ` Dave Chinner
2024-06-19 5:47 ` Christoph Hellwig
2024-06-19 13:44 ` Chuck Lever III
2024-06-19 21:16 ` NeilBrown
2024-06-19 0:42 ` Dave Chinner
2024-06-19 1:01 ` NeilBrown
2024-06-19 21:25 ` NeilBrown
2024-06-20 2:29 ` Dave Chinner [this message]
2024-06-20 10:18 ` Jeff Layton
2024-06-20 21:39 ` NeilBrown
2024-06-20 18:33 ` Chuck Lever III
2024-06-20 22:04 ` NeilBrown
2024-06-20 23:57 ` Trond Myklebust
2024-06-18 19:38 ` Jeff Layton
2024-06-18 23:12 ` NeilBrown
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ZnOUG2Nh80vTJXxe@dread.disaster.area \
--to=david@fromorbit.com \
--cc=chuck.lever@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neilb@suse.de \
--cc=trondmy@hammerspace.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox