From: Valeriy Glushkov <valeriy.glushkov at starwind.com>
To: spdk@lists.01.org
Subject: Re: [SPDK] A problem with SPDK 19.01 NVMeoF/RDMA target
Date: Tue, 26 Feb 2019 22:21:16 +0000 [thread overview]
Message-ID: <op.zxua1mvxns8zkg@sam8> (raw)
In-Reply-To: EA913ED399BBA34AA4EAC2EDC24CDD009C25E7B4@FMSMSX105.amr.corp.intel.com
[-- Attachment #1: Type: text/plain, Size: 4034 bytes --]
Hi Seth,
It seems that some problem is still present in the nvmf target built from
the recent SPDK trunk.
It crashes with dump under our performance test.
===========
# Starting SPDK v19.04-pre / DPDK 18.11.0 initialization...
[ DPDK EAL parameters: nvmf --no-shconf -c 0x1 --log-level=lib.eal:6
--base-virtaddr=0x200000000000 --file-prefix=spdk_pid131804 ]
EAL: No free hugepages reported in hugepages-2048kB
EAL: 2 hugepages of size 1073741824 reserved, but no mounted hugetlbfs
found for that size
app.c: 624:spdk_app_start: *NOTICE*: Total cores available: 1
reactor.c: 233:_spdk_reactor_run: *NOTICE*: Reactor started on core 0
rdma.c:2758:spdk_nvmf_rdma_poller_poll: *ERROR*: data=0x20002458d000
length=4096
rdma.c:2758:spdk_nvmf_rdma_poller_poll: *ERROR*: data=0x20002464b000
length=4096
rdma.c:2758:spdk_nvmf_rdma_poller_poll: *ERROR*: data=0x2000245ae000
length=4096
rdma.c:2786:spdk_nvmf_rdma_poller_poll: *ERROR*: data=0x20002457e000
length=4096
rdma.c:2758:spdk_nvmf_rdma_poller_poll: *ERROR*: data=0x20002457e000
length=4096
rdma.c:2786:spdk_nvmf_rdma_poller_poll: *ERROR*: data=0x2000245a8000
length=4096
nvmf_tgt: rdma.c:2789: spdk_nvmf_rdma_poller_poll: Assertion
`rdma_req->num_outstanding_data_wr > 0' failed.
^C
[1]+ Aborted (core dumped) ./nvmf_tgt -c ./nvmf.conf
=========
Do you need the dump file or some additional info?
--
Best regards,
Valeriy Glushkov
www.starwind.com
valeriy.glushkov(a)starwind.com
Howell, Seth <seth.howell(a)intel.com> писал(а) в своём письме Thu, 07 Feb
2019 20:18:28 +0200:
> Hi Sasha, Valeriy,
>
> With the help of Valeriy's logs I was able to get to the bottom of this.
> The root cause is that for NVMe-oF requests that don't transfer any
> data, such as keep_alive, we were not properly resetting the value of
> rdma_req->num_outstanding_data_wr between uses of that structure. All
> data carrying operations properly reset this value in
> spdk_nvmf_rdma_req_parse_sgl.
>
> My local repro steps look like this for anyone interested.
>
> Start the SPDK target,
> Submit a full queue depth worth of Smart log requests (sequentially is
> fine). A smaller number also works, but takes much longer.
> Wait for a while (This assumes you have keep alive enabled). Keep alive
> requests will reuse the rdma_req objects slowly incrementing the
> curr_send_depth on the admin qpair.
> Eventually the admin qpair will be unable to submit I/O.
>
> I was able to fix the issue locally with the following patch.
> https://review.gerrithub.io/#/c/spdk/spdk/+/443811/. Valeriy, please let
> me know if applying this patch also fixes it for you ( I am pretty sure
> that it will).
>
> Thank you for the bug report and for all of your help,
>
> Seth
>
> -----Original Message-----
> From: SPDK [mailto:spdk-bounces(a)lists.01.org] On Behalf Of Sasha
> Kotchubievsky
> Sent: Thursday, February 7, 2019 11:06 AM
> To: spdk(a)lists.01.org
> Subject: Re: [SPDK] A problem with SPDK 19.01 NVMeoF/RDMA target
>
> Hi,
>
> RNR value shouldn't affect NVMF. I just want to check if NVMF prepost
> enough receive requests. 19.10 introduced some new way for flow control
> and count number of send and receive work requests. Probably, NVMF
> doesn't pre-post enough requests.
>
> Which network do you use : IB or ROcE? What it is you HW and SW stack in
> host and in target sides? (OS, OFED/MOFED version, NIC type)
>
> I'd suggest to configure NVMF with big max queue depth, and in your test
> actually use a half of the value.
>
> On 2/7/2019 5:37 PM, Valeriy Glushkov wrote:
>> Hi Sasha,
>>
>> There is no IBV on the host side, it's Windows.
>> So we have no control over the RNR field.
>>
>> From a RDMA session's dump I can see that the initiator sets
>> infiniband.cm.req.rnrretrcount to 0x6.
>>
>> Could the RNR value be related to the problem we have with SPDK 19.01
>> NVMeoF target?
>>
next reply other threads:[~2019-02-26 22:21 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2019-02-26 22:21 Valeriy Glushkov [this message]
-- strict thread matches above, loose matches on Subject: below --
2019-02-28 22:07 [SPDK] A problem with SPDK 19.01 NVMeoF/RDMA target Valeriy Glushkov
2019-02-27 15:04 Howell, Seth
2019-02-27 15:00 Howell, Seth
2019-02-26 23:45 Lance Hartmann ORACLE
2019-02-07 18:18 Howell, Seth
2019-02-07 18:05 Sasha Kotchubievsky
2019-02-07 16:59 Howell, Seth
2019-02-07 15:37 Valeriy Glushkov
2019-02-07 14:31 Sasha Kotchubievsky
2019-02-07 4:50 Valeriy Glushkov
2019-02-06 23:06 Harris, James R
2019-02-06 23:03 Howell, Seth
2019-02-06 20:27 Howell, Seth
2019-02-06 18:15 Valeriy Glushkov
2019-02-06 16:54 Harris, James R
2019-02-06 16:03 Howell, Seth
2019-02-06 14:55 Harris, James R
2019-02-06 7:59 Valeriy Glushkov
2019-02-06 0:35 Harris, James R
2019-02-05 11:41 Valeriy Glushkov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=op.zxua1mvxns8zkg@sam8 \
--to=spdk@lists.01.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.