Linux-NVME Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: leonro@mellanox.com (Leon Romanovsky)
Subject: [PATCH v2 0/3] nvmet-rdma: SRQ per completion vector
Date: Thu, 16 Nov 2017 20:08:53 +0200	[thread overview]
Message-ID: <20171116180853.GN18825@mtr-leonro.local> (raw)
In-Reply-To: <1510852885-25519-1-git-send-email-maxg@mellanox.com>

On Thu, Nov 16, 2017@07:21:22PM +0200, Max Gurtovoy wrote:
> Since there is an active discussion regarding the CQ pool architecture, I decided to push
> this feature (maybe it can be pushed before CQ pool).

Max,

Thanks for CCing me, can you please repost the series and CC linux-rdma too?

>
> This is a new feature for NVMEoF RDMA target, that is intended to save resource allocation
> (by sharing them) and utilize the locality of completions to get the best performance with
> Shared Receive Queues (SRQs). We'll create a SRQ per completion vector (and not per device)
> using a new API (SRQ pool, added to this patchset too) and associate each created QP/CQ with
> an appropriate SRQ. This will also reduce the lock contention on the single SRQ per device
> (today's solution).
>
> My testing environment included 4 initiators (CX5, CX5, CX4, CX3) that were connected to 4
> subsystems (1 ns per sub) throw 2 ports (each initiator connected to unique subsystem
> backed in a different bull_blk device) using a switch to the NVMEoF target (CX5).
> I used RoCE link layer.
>
> Configuration:
>  - Irqbalancer stopped on each server
>  - set_irq_affinity.sh on each interface
>  - 2 initiators run traffic throw port 1
>  - 2 initiators run traffic throw port 2
>  - On initiator set register_always=N
>  - Fio with 12 jobs, iodepth 128
>
> Memory consumption calculation for recv buffers (target):
>  - Multiple SRQ: SRQ_size * comp_num * ib_devs_num * inline_buffer_size
>  - Single SRQ: SRQ_size * 1 * ib_devs_num * inline_buffer_size
>  - MQ: RQ_size * CPU_num * ctrl_num * inline_buffer_size
>
> Cases:
>  1. Multiple SRQ with 1024 entries:
>     - Mem = 1024 * 24 * 2 * 4k = 192MiB (Constant number - not depend on initiators number)
>  2. Multiple SRQ with 256 entries:
>     - Mem = 256 * 24 * 2 * 4k = 48MiB (Constant number - not depend on initiators number)
>  3. MQ:
>     - Mem = 256 * 24 * 8 * 4k = 192MiB (Mem grows for every new created ctrl)
>  4. Single SRQ (current SRQ implementation):
>     - Mem = 4096 * 1 * 2 * 4k = 32MiB (Constant number - not depend on initiators number)
>
> results:
>
> BS    1.read (target CPU)   2.read (target CPU)    3.read (target CPU)   4.read (target CPU)
> ---  --------------------- --------------------- --------------------- ----------------------
> 1k     5.88M (80%)            5.45M (72%)            6.77M (91%)          2.2M (72%)
>
> 2k     3.56M (65%)            3.45M (59%)            3.72M (64%)          2.12M (59%)
>
> 4k     1.8M (33%)             1.87M (32%)            1.88M (32%)          1.59M (34%)
>
> BS    1.write (target CPU)   2.write (target CPU) 3.write (target CPU)   4.write (target CPU)
> ---  --------------------- --------------------- --------------------- ----------------------
> 1k     5.42M (63%)            5.14M (55%)            7.75M (82%)          2.14M (74%)
>
> 2k     4.15M (56%)            4.14M (51%)            4.16M (52%)          2.08M (73%)
>
> 4k     2.17M (28%)            2.17M (27%)            2.16M (28%)          1.62M (24%)
>
>
> We can see the perf improvement between Case 2 and Case 4 (same order of resource).
> We can see the benefit in resource consumption (mem and CPU) with a small perf loss
> between cases 2 and 3.
> There is still an open question between the perf differance for 1k between Case 1 and
> Case 3, but I guess we can investigate and improve it incrementaly.
>
> Thanks to Idan Burstein and Oren Duer for suggesting this nice feature.
>
> Changes from V1:
>  - Added SRQ pool per protection domain for IB/core
>  - Fixed few comments from Christoph and Sagi
>
> Max Gurtovoy (3):
>   IB/core: add a simple SRQ pool per PD
>   nvmet-rdma: use srq pointer in rdma_cmd
>   nvmet-rdma: use SRQ per completion vector
>
>  drivers/infiniband/core/Makefile   |   2 +-
>  drivers/infiniband/core/srq_pool.c | 106 +++++++++++++++++++++
>  drivers/infiniband/core/verbs.c    |   4 +
>  drivers/nvme/target/rdma.c         | 190 +++++++++++++++++++++++++++----------
>  include/rdma/ib_verbs.h            |   5 +
>  include/rdma/srq_pool.h            |  46 +++++++++
>  6 files changed, 301 insertions(+), 52 deletions(-)
>  create mode 100644 drivers/infiniband/core/srq_pool.c
>  create mode 100644 include/rdma/srq_pool.h
>
> --
> 1.8.3.1
>

  parent reply	other threads:[~2017-11-16 18:08 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2017-11-16 17:21 [PATCH v2 0/3] nvmet-rdma: SRQ per completion vector Max Gurtovoy
2017-11-16 17:21 ` [PATCH 1/3] IB/core: add a simple SRQ pool per PD Max Gurtovoy
2017-11-16 18:10   ` Leon Romanovsky
2017-11-17 19:12     ` Max Gurtovoy
2017-11-18  7:36       ` Leon Romanovsky
2017-11-16 18:32   ` Sagi Grimberg
2017-11-17 19:17     ` Max Gurtovoy
2017-11-16 17:21 ` [PATCH 2/3] nvmet-rdma: use srq pointer in rdma_cmd Max Gurtovoy
2017-11-16 17:26   ` [Suspected-Phishing][PATCH " Max Gurtovoy
2017-11-16 18:37   ` [PATCH " Sagi Grimberg
2017-11-16 17:21 ` [PATCH 3/3] nvmet-rdma: use SRQ per completion vector Max Gurtovoy
2017-11-16 18:08 ` Leon Romanovsky [this message]
2017-11-17 19:18   ` [PATCH v2 0/3] nvmet-rdma: " Max Gurtovoy
2017-11-16 18:36 ` Sagi Grimberg
2017-11-17 19:32   ` Max Gurtovoy
2017-11-18 12:52     ` Leon Romanovsky
2017-11-18 13:57       ` Max Gurtovoy
2017-11-18 14:40         ` Leon Romanovsky
2017-11-18 21:29           ` Max Gurtovoy
2017-11-20 11:00             ` Sagi Grimberg
2017-11-20 11:34               ` Leon Romanovsky

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20171116180853.GN18825@mtr-leonro.local \
    --to=leonro@mellanox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox