The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Tom Tucker <tom@opengridcomputing.com>
To: Jeff Layton <jlayton@redhat.com>
Cc: linux-kernel@vger.kernel.org, linux-nfs@vger.kernel.org,
	bfields@fieldses.org
Subject: Re: [PATCH 0/3] have pooled sunrpc services make more intelligent allocations
Date: Tue, 03 Jun 2008 13:37:25 -0500	[thread overview]
Message-ID: <1212518245.29133.25.camel@trinity.ogc.int> (raw)
In-Reply-To: <20080603134257.1de3a1c3@tleilax.poochiereds.net>


On Tue, 2008-06-03 at 13:42 -0400, Jeff Layton wrote:
> On Tue, 03 Jun 2008 11:53:42 -0500
> Tom Tucker <tom@opengridcomputing.com> wrote:
> 
> > Jeff:
> > 
> > This brings up an interesting issue with the RDMA transport and
> > RDMA_READ. RDMA_READ is submitted as part of fetching an RPC from the
> > client (e.g. NFS_WRITE). The xpo_recvfrom function doesn't block waiting
> > for the RDMA_READ to complete, but rather queues the RPC for subsequent
> > processing when the I/O completes and returns 0. 
> > 
> > I can use these new services to allocate CPU local pages for this I/O.
> > So far, so good. However, when the I/O completes, and the transport is
> > rescheduled for subsequent RPC completion processing, the pool/CPU that
> > is elected doesn't have any affinity for the CPU on which the I/O was
> > initially submitted. I think this means that the svc_process/reply steps
> > may occur on a CPU far away from the memory in which the data resides.
> > 
> > Am I making sense here? If so, any thoughts on what could/should be
> > done?
> > 
> > Thanks,
> > Tom
> > 
> 
> I confess I didn't think hard about the RDMA case here (and haven't
> been paying as much attention as I probably should to the design of
> it). So take my thoughts with a large chunk of salt...
> 
> On a NUMA box, the pages have to live _somewhere_ and some CPUs will be
> closer to them than others. If we're concerned about making sure that
> the post-RDMA_READ processing is done on a CPU close to the memory,
> then we don't have much choice but to try to make sure that this
> processing is only done on CPUs that are close to that memory.
> 
> Assuming that this post-processing is done by nfsd, I suppose we'd need
> to tag the post-RDMA_READ RPC with a poolid or something and make sure
> that only nfsds running on CPUs close to the memory pick it up. Perhaps
> there could be a per-pool queue for these RPC's or something...
> 
> Either way, the big question is whether that will be a net win or loss
> for throughput. i.e. are we better off waiting for the right nfsd to
> become available or allowing the first nfsd that becomes available to
> make the crosscalls needed to do the RPC? It's hard to say...

Not only that, but it would lead to more disorder in the RPC processing
which might kill write-behind.

> 
> In the near term, I doubt this patchset will harm the RDMA case. 

Agreed. 

> After
> all, the distribution of memory allocations is pretty lumpy now. On
> a NUMA box with RDMA you're probably doing a lot of crosscalls with
> the current code.

Probably no worse than the socket's transport since the skbuf's aren't
necessarily allocated on the CPU calling svc_recv.

> 




  reply	other threads:[~2008-06-03 18:32 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2008-06-03 11:16 [PATCH 0/3] have pooled sunrpc services make more intelligent allocations Jeff Layton
2008-06-03 16:53 ` Tom Tucker
2008-06-03 17:42   ` Jeff Layton
2008-06-03 18:37     ` Tom Tucker [this message]
2008-06-04 11:53       ` Jeff Layton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1212518245.29133.25.camel@trinity.ogc.int \
    --to=tom@opengridcomputing.com \
    --cc=bfields@fieldses.org \
    --cc=jlayton@redhat.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox