From: Sagi Grimberg <sagi@grimberg.me>
To: Daniel Wagner <dwagner@suse.de>,
"Knight, Frederick" <Frederick.Knight@netapp.com>
Cc: "linux-nvme@lists.infradead.org" <linux-nvme@lists.infradead.org>,
Shinichiro Kawasaki <shinichiro.kawasaki@wdc.com>,
"hare@suse.de" <hare@suse.de>
Subject: Re: [PATCH v2] nvmet: force reconnect when number of queue changes
Date: Wed, 28 Sep 2022 17:23:06 +0300 [thread overview]
Message-ID: <b121c092-cecc-b891-be3c-b32c2a3e611d@grimberg.me> (raw)
In-Reply-To: <20220928135053.s6h6xr3llhp6riwr@carbon.lan>
>> Would this be a case for a new AEN - controller configuration changed?
>> I'm also wondering exactly what changed in the controller? It can't
>> be the "Number of Queues" feature (that can't change - The controller
>> shall not change the value allocated between resets.). Is it the MQES
>> field in the CAP property that changes (queue size)?
>>
>> We already have change reporting for: Namespace attribute, Predictable
>> Latency, LBA status, EG aggregate, Zone descriptor, Discovery Log,
>> Reservations. We've questioned whether we need a Controller Attribute
>> Changed.
>>
>> Would this be a case for an exception? Does the DNR bit apply only to
>> commands sent on queues that already exist (i.e., NOT the connect
>> command since that command is actually creating the queue)? FWIW - I
>> don't like exceptions.
>>
>> Can you elaborate on exactly what is changing?
>
> The background story is, that we have a setup with two targets (*) and
> the host is connected two both of them (HA setup). Both server run the
> same software version. The host is told via Number of Queues (Feature
> Identifier 07h) how many queues are supported (N queues).
>
> Now, a software upgrade is started which takes first one server offline,
> updates it and brings it online again. Then the same process with the
> second server.
>
> In the meantime the host tries to reconnect. Eventually, the reconnect
> is successful but the Number of Queues (Feature Identifier 07h) has
> changed to a smaller value, e.g N-2 queues.
>
> My test case here is trying to replicated this scenario but just with
> one target. Hence the problem how to notify the host that it should
> reconnect. As you mentioned this is not to supposed to change as long a
> connection is established.
>
> My understanding is that the current nvme target implementation in Linux
> doesn't really support this HA setup scenario hence my attempt to get it
> flying with one target. The DNR bit comes into play because I was toying
> with removing the subsystem from the port, changing the number of queues
> and re-adding the subsystem to the port.
>
> This resulted in any request posted from the host seeing the DNR
> bit. The second attempt here was to delete the controller to force a
> reconnect. I agree it's also not really the right thing to do.
>
> As far I can tell, what's is missing from a testing point of view is the
> ability to fail requests without the DNR bit set or the ability to tell
> the host to reconnect. Obviously, an AEN would be nice for this but I
> don't know if this is reason enough to extend the spec.
Looking into the code, its the connect that fails on invalid parameter
with a DNR, because the host is attempting to connect to a subsystems
that does not exist on the port (because it was taken offline for
maintenance reasons).
So I guess it is valid to allow queue change without removing it from
the port, but that does not change the fundamental question on DNR.
If the host sees a DNR error on connect, my interpretation is that the
host should not retry the connect command itself, but it shouldn't imply
anything on tearing down the controller and giving up on it completely,
forever.
next prev parent reply other threads:[~2022-09-28 14:23 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2022-09-27 14:31 [PATCH v2] nvmet: force reconnect when number of queue changes Daniel Wagner
2022-09-27 15:01 ` Daniel Wagner
2022-09-28 6:55 ` Sagi Grimberg
2022-09-28 7:48 ` Daniel Wagner
2022-09-28 8:31 ` Sagi Grimberg
2022-09-28 9:02 ` Daniel Wagner
2022-09-28 12:39 ` Knight, Frederick
2022-09-28 13:50 ` Daniel Wagner
2022-09-28 14:23 ` Sagi Grimberg [this message]
2022-09-28 15:21 ` Hannes Reinecke
2022-09-28 16:01 ` Knight, Frederick
2022-10-05 18:15 ` Daniel Wagner
2022-10-06 11:37 ` Sagi Grimberg
2022-10-06 20:15 ` James Smart
2022-10-06 20:54 ` Knight, Frederick
2022-09-28 16:01 ` Knight, Frederick
2022-09-28 15:01 ` John Meneghini
2022-09-28 15:26 ` Daniel Wagner
2022-09-28 18:02 ` Knight, Frederick
2022-09-29 2:14 ` John Meneghini
2022-09-29 3:04 ` Knight, Frederick
2022-09-30 7:03 ` Hannes Reinecke
2022-09-30 6:32 ` Hannes Reinecke
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b121c092-cecc-b891-be3c-b32c2a3e611d@grimberg.me \
--to=sagi@grimberg.me \
--cc=Frederick.Knight@netapp.com \
--cc=dwagner@suse.de \
--cc=hare@suse.de \
--cc=linux-nvme@lists.infradead.org \
--cc=shinichiro.kawasaki@wdc.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox