Hi, I'm using vpp 19.04.2 and spdk 19.07. I have a couple questions about the architecture of the session API with nvmeotcp. I attached my perf results for a 4k random read. I have 3 cores for vpp workers with 3 RSS queues and the other 5 cores are for spdk. The memcpy is the bottleneck taking 31% cpu usage on the spdk threads when it writes the completion into shared memory. The vpp thread memcpy is 54% when it pulls the data out for tx. You can flip those results around for a random write. L3 cache miss is bad at 60+%. With RDMA it's ~10%. This gives rather poor performance. Memory is more of a bottleneck than cpu on my ARM system, so the spdk design worked out well being zero copy. It would be better if the copying in and out of the shmem were avoided though I'm not sure it's feasible and need to study the code further. Are there any plans to change the session API or the way it uses memory? Spdk and vpp both use dpdk independently with different memory regions (--file-prefix). But the shared memory in the session API uses a libc malloc'd buffer. Would using a dpdk mempool be better for cache efficiency and synchronization with other dpdk memory usage? Any thoughts on shared dpdk mempools between vpp and spdk instead of copying into the session API's shared memory? It looks like there are 2 copies for 1 read, the nvme completion into shmem, and the vpp copy out (haven't looked where session_tx_fifo_peek_and_snd is copying to). Even getting rid of one of the memcpy's could help. We are looking into possible optimizations so any guidance is appreciated since we want any changes merged into the mainline. Thanks, Jon