Linux-NVME Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: John Meneghini <jmeneghi@redhat.com>
To: Jesse Taube <jtaubepe@redhat.com>, linux-block@vger.kernel.org
Cc: linux-nvme@lists.infradead.org, shinichiro.kawasaki@wdc.com,
	Daniel Wagner <dwagner@suse.de>
Subject: Re: [PATCH v2 0/2] Test multipath and marginal ports
Date: Thu, 20 Aug 2026 10:09:28 -0400	[thread overview]
Message-ID: <94eaddfb-36a0-4ac6-9ea4-6eaad46aaf61@redhat.com> (raw)
In-Reply-To: <0d3cb83e-a4b6-414d-8756-fafd4a8238bc@redhat.com>

Further analysis shows that this is a preexisting issue with test/nvme/057.  So I will be trying to fix that test before I combine it with test/nvme/070.

target-vm:blktests(fpin_tests10) > sudo NVMET_TRTYPES=fc ./check tests/nvme/057
nvme/057 (tr=fc) (test nvme fabrics controller ANA failover during I/O) [passed]
     runtime  25.421s  ...  25.411s
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 8210 vs 8210
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 49169 vs 49169
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 45075 vs 45075
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 16404 vs 16404
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 28693 vs 28693
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 32790 vs 32790
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 28695 vs 28695
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: transport association event: transport detected io error
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: resetting controller
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: create association : host wwpn 0x20001100aa000001  rport wwpn 0x20001100ab000003: NQN "blktests-subsystem-1"
[Thu Aug 20 10:06:04 2026] (NULL device *): {2:1} Association created
[Thu Aug 20 10:06:04 2026] nvmet: Created nvm controller 1 for subsystem blktests-subsystem-1 for NQN nqn.2014-08.org.nvmexpress:uuid:0f01fb42-9f7f-4856-b0b3-51e60b8de349.
[Thu Aug 20 10:06:04 2026] nvme nvme5: NVME-FC{3}: controller connect complete
[Thu Aug 20 10:06:04 2026] (NULL device *): {2:0} Association deleted
[Thu Aug 20 10:06:04 2026] (NULL device *): {2:0} Association freed
[Thu Aug 20 10:06:04 2026] (NULL device *): Disconnect LS failed: No Association
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 45084 vs 45084
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 16413 vs 16413
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 8222 vs 8222
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: transport association event: transport detected io error
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: resetting controller
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 45087 vs 45087
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 24608 vs 24608
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 4113 vs 4113
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 40978 vs 40978
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 53267 vs 53267
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 24596 vs 24596
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 53269 vs 53269
[Thu Aug 20 10:06:16 2026] nvme nvme2: NVME-FC{0}: io failed due to bad NVMe_ERSP: iu len 8, xfr len 4096 vs 0, status code 0, cmdid 8214 vs 8214
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O
[Thu Aug 20 10:06:16 2026] block nvme2n1: no usable path - requeuing I/O

On 8/19/26 17:49, John Meneghini wrote:
> This is a great improvement and we now have a functioning blkstest 070 to test the FPIN LI kernel patches.
> 
> The problem is: I don't think we are done yet.  In my own private testing and development with these patches I've been working on a next-version test that combines the ANA states from test/nvme/057 with test/nvme/070. So I am working on a test 071.
> 
> The good news is: everything now works with test/nvme/070.  The bad news is: the kernel patches are not done.
> 
> What I've found is: a long as the ANA states are all optimized or non-optimized everything works.  However, once we throw in an inaccessible state to the mix, we run into serious problems.
> 
> At this point in development I don't care about the test failures in my test/nvme/071 script. I expect the script to bug out because it doesn't understand the inaccessible state. The test sill continues flipping rports in and out of the marginal state and keeps going. That's what it is designed to do. That means it is testing all of the code paths in the kernel patches.
> 
> The problem is: when turning marginal paths on and off with controllers that are in the inaccessible state, the path selection algorithm in the kernel fails and we end up with the following:
> 
> [Wed Aug 19 16:25:42 2026] nvme_ns_head_submit_bio: 6 callbacks suppressed
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> [Wed Aug 19 16:25:42 2026] block nvme2n1: no usable path - requeuing I/O
> 
> At this point the fio jobs are still running but there is no progress.  This means no path was found and ALL of the IOs got re-queued. And if we flip the marginal state off on all of the controllers IO continues to be hung.  The IO scheduler is hung and there is no possibility of getting it restarted again.
> 
> So this is a really serious bug in the kernel patches and we can't ship this stuff until we fix the problem.
> 
> This IO re-requeing problem should NEVER happen - no matter what the state of the marginal paths.
> 
> So I can recommend that Shinichiro test these patches with the current upstream kernel patches:
> 
>    https://lore.kernel.org/linux-nvme/20260812181300.3712426-1-jtaubepe@redhat.com/
> But there will be another version of Kernel patches and a V3 of this patch set will be forth coming.
> 
> John A. Meneghini
> Senior Principal Platform Storage Engineer
> RHEL SST - Platform Storage Group
> jmeneghi@redhat.com
> 
> On 8/19/26 16:04, Jesse Taube wrote:
>> Tests for the upcoming nvme-fc: FPIN link integrity handling set.
>> It tests for various multipath and marginal port
>> scenarios, while confirming the port usage and state. The test is
>> intended to emulate receiving an FPIN event in a multipath environment.
>>
>> Link: https://bugzilla.kernel.org/show_bug.cgi?id=220329
>> Link: https://github.com/linux-blktests/blktests/pull/264
>>
>> Jesse Taube (2):
>>    nvme: Add _setup_nvmet_port_marginal
>>    nvme/070: Test multipath and marginal ports
>>
>>   common/nvme        |  31 +++
>>   tests/nvme/070     | 613 +++++++++++++++++++++++++++++++++++++++++++++
>>   tests/nvme/070.out |  43 ++++
>>   3 files changed, 687 insertions(+)
>>   create mode 100755 tests/nvme/070
>>   create mode 100644 tests/nvme/070.out
>>
> 



  reply	other threads:[~2026-08-20 14:09 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-19 20:04 [PATCH v2 0/2] Test multipath and marginal ports Jesse Taube
2026-08-19 20:04 ` [PATCH v2 1/2] nvme: Add _setup_nvmet_port_marginal Jesse Taube
2026-08-19 20:04 ` [PATCH v2 2/2] nvme/070: Test multipath and marginal ports Jesse Taube
2026-08-19 21:49 ` [PATCH v2 0/2] " John Meneghini
2026-08-20 14:09   ` John Meneghini [this message]
2026-08-23 12:24   ` Shin'ichiro Kawasaki

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=94eaddfb-36a0-4ac6-9ea4-6eaad46aaf61@redhat.com \
    --to=jmeneghi@redhat.com \
    --cc=dwagner@suse.de \
    --cc=jtaubepe@redhat.com \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-nvme@lists.infradead.org \
    --cc=shinichiro.kawasaki@wdc.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox