From mboxrd@z Thu Jan 1 00:00:00 1970 Content-Type: multipart/mixed; boundary="===============8371272858674945391==" MIME-Version: 1.0 From: Walker, Benjamin Subject: Re: [SPDK] spdk/vpp performance Date: Mon, 16 Sep 2019 19:53:34 +0000 Message-ID: <97254b5b1571de4f7664a95dad4bad331f18f689.camel@intel.com> In-Reply-To: ec19421d1822d9cf2f0cfeed76ab0adc@mail.gmail.com List-ID: To: spdk@lists.01.org --===============8371272858674945391== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable On Tue, 2019-09-10 at 12:42 -0700, Jonathan Richardson via SPDK wrote: > Hi, > = > I'm using vpp 19.04.2 and spdk 19.07. I have a couple questions about the > architecture of the session API with nvmeotcp. I attached my perf results > for a 4k random read. I have 3 cores for vpp workers with 3 RSS queues and > the other 5 cores are for spdk. The memcpy is the bottleneck taking 31% > cpu usage on the spdk threads when it writes the completion into shared > memory. The vpp thread memcpy is 54% when it pulls the data out for tx. > You can flip those results around for a random write. L3 cache miss is bad > at 60+%. With RDMA it's ~10%. This gives rather poor performance. Memory > is more of a bottleneck than cpu on my ARM system, so the spdk design > worked out well being zero copy. It would be better if the copying in and > out of the shmem were avoided though I'm not sure it's feasible and need > to study the code further. > = > Are there any plans to change the session API or the way it uses memory? I've been investigating a bit here and I'm struggling to figure out specifi= cally what the optimal way to use the session API would be here - it's undocument= ed as far as I can tell. I wish I had an answer, but I don't. Certainly if we cou= ld get some experts on VPP to weight in it would be awesome. > = > Spdk and vpp both use dpdk independently with different memory regions > (--file-prefix). But the shared memory in the session API uses a libc > malloc'd buffer. Would using a dpdk mempool be better for cache efficiency > and synchronization with other dpdk memory usage? If the DMA operations could occur directly to this memory, yes. If due to s= ome other architecture reason it still copies into and out of it, then it won't= help to make it a memory pool. > = > Any thoughts on shared dpdk mempools between vpp and spdk instead of > copying into the session API's shared memory? It looks like there are 2 > copies for 1 read, the nvme completion into shmem, and the vpp copy out > (haven't looked where session_tx_fifo_peek_and_snd is copying to). Even > getting rid of one of the memcpy's could help. > = > We are looking into possible optimizations so any guidance is appreciated > since we want any changes merged into the mainline. The guidance we've received from VPP on this is that the copies are unavoid= able via the session or vcl APIs. To avoid the copies, we'd have to make SPDK a = VPP plugin and run within their framework. I haven't entirely comprehended what= that would mean for us. SPDK is designed to integrate reasonably well into other cooperative multi-tasking frameworks. For example, see here where Jim integ= rated SPDK into Seastar: https://review.gerrithub.io/c/spdk/spdk/+/466629 In order to integrate into another framework, we just need to call spdk_thread_lib_init and pass in a function pointer that can schedule our lightweight threads onto whatever event loop the framework has. I think that means creating "input" nodes for each lightweight thread so that they get p= olled on each pass through the loop, and then re-implementing our sock abstractio= n so that it just passes any data to the "next node" which would be the TCP processing node. But there's obviously a huge amount of detail to be filled= in beyond what I've just said, and I have no idea currently how to do it. > = > Thanks, > Jon > _______________________________________________ > SPDK mailing list > SPDK(a)lists.01.org > https://lists.01.org/mailman/listinfo/spdk --===============8371272858674945391==--