From mboxrd@z Thu Jan 1 00:00:00 1970 From: Karan Vohra Subject: Questions around multipath failover and no_path_retry Date: Sun, 11 Mar 2018 16:47:17 +0000 Message-ID: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1428333661117937983==" Return-path: Content-Language: en-US List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: dm-devel-bounces@redhat.com Errors-To: dm-devel-bounces@redhat.com To: "dm-devel@redhat.com" List-Id: dm-devel.ids --===============1428333661117937983== Content-Language: en-US Content-Type: multipart/alternative; boundary="_000_SN1PR20MB047927660CFE0409EDE91C95C9DC0SN1PR20MB0479namp_" --_000_SN1PR20MB047927660CFE0409EDE91C95C9DC0SN1PR20MB0479namp_ Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: quoted-printable Hi Folks, Let us assume, there are 2 paths within the path group which dm-multipath i= s sending the I/Os in round-robin fashion. Each of these paths are identifi= ed as unique block device(s) such as /dev/sdb and /dev/sdc. Let us say some I/Os are sent over to the path /dev/sdb and either the requ= ests time out or there is a failure on that path, what happens to those I/O= s? Are they sent over to the other path - /dev/sdc or does dm-multipath wai= ts for /dev/sdb to come back online and only sends I/O to /dev/sdb? One of = the reasons we are concerned about the above scenario is- let us say there = is a write I/O W1 which is routed to /dev/sdb and then there is a failure. = There was a write I/O W2 which wrote at the same block via /dev/sdc. Now if= multipath sends W1 through /dev/sdc, W2 gets overwritten by W1. The expect= ation was that W2 happens after W1 and should overwrite W1 but the result i= s opposite. Situations like these can cause data inconsistency and corrupti= on.We were thinking of using no_path_retry configuration to be set to queue= to make sure that the I/Os supposed to be going to path1 never make it to = path2. But the question is that would not that cause unexpected behavior in= application layer? Let us say there are I/O Requests R1, R2, R3 and so on.= . R1 is going to Path1, R2 is going to Path2 and so on. If Path1 dies for s= ome reason, with the setting of no_path_retry to queue, queueing will not s= top until the path is fixed so does not that mean that R1, R3,R5 ... will n= ot make it to block device until the path is fixed? Would it not cause fail= ures if the issue persists for seconds? What about the size of queue? Is th= ere any danger of queue getting overloaded?Any pointers or references would= be of great help. Thanks! Karan Get Outlook for Android --_000_SN1PR20MB047927660CFE0409EDE91C95C9DC0SN1PR20MB0479namp_ Content-Type: text/html; charset="us-ascii" Content-Transfer-Encoding: quoted-printable
Hi Folks,

Let us assume, there are 2 paths within the path group which dm-multipath i= s sending the I/Os in round-robin fashion. Each of these paths are identifi= ed as unique block device(s) such as /dev/sdb and /dev/sdc. 

Let us say some I/Os are sent over to the path /dev/sdb and either the requ= ests time out or there is a failure on that path, what happens to those I/O= s? Are they sent over to the other path - /dev/sdc or does dm-multipath wai= ts for /dev/sdb to come back online and only sends I/O to /dev/sdb? One of the reasons we are concerned about = the above scenario is- let us say there is a write I/O W1 which is routed t= o /dev/sdb and then there is a failure. There was a write I/O W2 which wrot= e at the same block via /dev/sdc. Now if multipath sends W1 through /dev/sdc, W2 gets overwritten by W1. The exp= ectation was that W2 happens after W1 and should overwrite W1 but the resul= t is opposite. Situations like these can cause data inconsistency and corru= ption.We were thinking of using no_path_retry configuration to be set to queue to make sure that the I/Os = supposed to be going to path1 never make it to path2. But the question is t= hat would not that cause unexpected behavior in application layer? Let us s= ay there are I/O Requests R1, R2, R3 and so on.. R1 is going to Path1, R2 is going to Path2 and so on. If Pa= th1 dies for some reason, with the setting of no_path_retry to queue, queue= ing will not stop until the path is fixed so does not that mean that R1, R3= ,R5 ... will not make it to block device until the path is fixed? Would it not cause failures if the issue p= ersists for seconds? What about the size of queue? Is there any danger of q= ueue getting overloaded?Any pointers or references would be of great help.<= br>
Thanks!
Karan


--_000_SN1PR20MB047927660CFE0409EDE91C95C9DC0SN1PR20MB0479namp_-- --===============1428333661117937983== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline --===============1428333661117937983==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: Martin Wilck Subject: Re: Questions around multipath failover and no_path_retry Date: Mon, 19 Mar 2018 23:48:43 +0100 Message-ID: <1521499723.3798.153.camel@suse.com> References: Mime-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: base64 Return-path: In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: dm-devel-bounces@redhat.com Errors-To: dm-devel-bounces@redhat.com To: Karan Vohra , "dm-devel@redhat.com" List-Id: dm-devel.ids T24gU3VuLCAyMDE4LTAzLTExIGF0IDE2OjQ3ICswMDAwLCBLYXJhbiBWb2hyYSB3cm90ZToKPiBI aSBGb2xrcywKPiAKPiBMZXQgdXMgYXNzdW1lLCB0aGVyZSBhcmUgMiBwYXRocyB3aXRoaW4gdGhl IHBhdGggZ3JvdXAgd2hpY2ggZG0tCj4gbXVsdGlwYXRoIGlzIHNlbmRpbmcgdGhlIEkvT3MgaW4g cm91bmQtcm9iaW4gZmFzaGlvbi4gRWFjaCBvZiB0aGVzZQo+IHBhdGhzIGFyZSBpZGVudGlmaWVk IGFzIHVuaXF1ZSBibG9jayBkZXZpY2Uocykgc3VjaCBhcyAvZGV2L3NkYiBhbmQKPiAvZGV2L3Nk Yy4gCj4gCj4gTGV0IHVzIHNheSBzb21lIEkvT3MgYXJlIHNlbnQgb3ZlciB0byB0aGUgcGF0aCAv ZGV2L3NkYiBhbmQgZWl0aGVyCj4gdGhlIHJlcXVlc3RzIHRpbWUgb3V0IG9yIHRoZXJlIGlzIGEg ZmFpbHVyZSBvbiB0aGF0IHBhdGgsIHdoYXQKPiBoYXBwZW5zIHRvIHRob3NlIEkvT3M/IEFyZSB0 aGV5IHNlbnQgb3ZlciB0byB0aGUgb3RoZXIgcGF0aCAtCj4gL2Rldi9zZGMgb3IgZG9lcyBkbS1t dWx0aXBhdGggd2FpdHMgZm9yIC9kZXYvc2RiIHRvIGNvbWUgYmFjayBvbmxpbmUKPiBhbmQgb25s eSBzZW5kcyBJL08gdG8gL2Rldi9zZGI/CgpUaGUgSS9PIGlzIHNlbnQgdG8gb3RoZXJyIHBhdGhz IChzZGMgaW4geW91ciBleGFtcGxlKSB3aGVuIHRoZSBsb3dlcgpsYXllciAoZS5nLiBTQ1NJKSBp bmRpY2F0ZXMgcGF0aCBmYWlsdXJlIGZvciBzZGIuIFRoYXQncyB0aGUgdmVyeSBwb2ludApvZiBt dWx0aXBhdGhpbmcuCgo+ICBPbmUgb2YgdGhlIHJlYXNvbnMgd2UgYXJlIGNvbmNlcm5lZCBhYm91 dCB0aGUgYWJvdmUgc2NlbmFyaW8gaXMtIGxldAo+IHVzIHNheSB0aGVyZSBpcyBhIHdyaXRlIEkv TyBXMSB3aGljaCBpcyByb3V0ZWQgdG8gL2Rldi9zZGIgYW5kIHRoZW4KPiB0aGVyZSBpcyBhIGZh aWx1cmUuIFRoZXJlIHdhcyBhIHdyaXRlIEkvTyBXMiB3aGljaCB3cm90ZSBhdCB0aGUgc2FtZQo+ IGJsb2NrIHZpYSAvZGV2L3NkYy4gTm93IGlmIG11bHRpcGF0aCBzZW5kcyBXMSB0aHJvdWdoIC9k ZXYvc2RjLCBXMgo+IGdldHMgb3ZlcndyaXR0ZW4gYnkgVzEuIFRoZSBleHBlY3RhdGlvbiB3YXMg dGhhdCBXMiBoYXBwZW5zIGFmdGVyIFcxCj4gYW5kIHNob3VsZCBvdmVyd3JpdGUgVzEgYnV0IHRo ZSByZXN1bHQgaXMgb3Bwb3NpdGUuIAoKSWYgeW91IHNlbmQgdHdvIHdyaXRlIElPcyB0byB0aGUg c2FtZSBzZWN0b3IgYXQgdGhlIHNhbWUgdGltZSwgeW91CmNhbid0IGJlIHN1cmUgd2hpY2ggb25l IGFycml2ZXMgZmlyc3QuIFRoYXQncyBub3Qgc3BlY2lmaWMgdG8KbXVsdGlwYXRoLiBJZiB5b3Ug d2FudCB0byBndWFyYW50ZWUgb3JkZXJpbmcsIHlvdSBoYXZlIHRvIGZsdXNoIFcxIAp1c2luZyBl LmcuIGZkYXRhc3luYygpIGJlZm9yZSBzZW5kaW5nIFcyLiBUaGUgZmx1c2ggY29tbWFuZCB3b24n dApyZXR1cm4gYmVmb3JlIFcxIGlzIHdyaXR0ZW4gdG8gZGlzay4KCj4gU2l0dWF0aW9ucyBsaWtl IHRoZXNlIGNhbiBjYXVzZSBkYXRhIGluY29uc2lzdGVuY3kgYW5kIGNvcnJ1cHRpb24uV2UKPiB3 ZXJlIHRoaW5raW5nIG9mIHVzaW5nIG5vX3BhdGhfcmV0cnkgY29uZmlndXJhdGlvbiB0byBiZSBz ZXQgdG8gcXVldWUKPiB0byBtYWtlIHN1cmUgdGhhdCB0aGUgSS9PcyBzdXBwb3NlZCB0byBiZSBn b2luZyB0byBwYXRoMSBuZXZlciBtYWtlCj4gaXQgdG8gcGF0aDIuCgpUaGF0IHdvbid0IHdvcmsu IEFzIHRoZSBuYW1lIG9mIHRoZSBvcHRpb24gc3VnZ2Vlc3RzLCAibm9fcGF0aF9yZXRyeSIKb25s eSBhZmZlY3RzIHRoZSBiZWhhdmlvciBpZiB0aGVyZSdzIF9ub18gaGVhbHRoeSBwYXRoIGxlZnQu Cgo+ICBCdXQgdGhlIHF1ZXN0aW9uIGlzIHRoYXQgd291bGQgbm90IHRoYXQgY2F1c2UgdW5leHBl Y3RlZCBiZWhhdmlvciBpbgo+IGFwcGxpY2F0aW9uIGxheWVyPyBMZXQgdXMgc2F5IHRoZXJlIGFy ZSBJL08gUmVxdWVzdHMgUjEsIFIyLCBSMyBhbmQKPiBzbyBvbi4uIFIxIGlzIGdvaW5nIHRvIFBh dGgxLCBSMiBpcyBnb2luZyB0byBQYXRoMiBhbmQgc28gb24uIElmCj4gUGF0aDEgZGllcyBmb3Ig c29tZSByZWFzb24sIHdpdGggdGhlIHNldHRpbmcgb2Ygbm9fcGF0aF9yZXRyeSB0bwo+IHF1ZXVl LCBxdWV1ZWluZyB3aWxsIG5vdCBzdG9wIHVudGlsIHRoZSBwYXRoIGlzIGZpeGVkIHNvIGRvZXMg bm90Cj4gdGhhdCBtZWFuIHRoYXQgUjEsIFIzLFI1IC4uLiB3aWxsIG5vdCBtYWtlIGl0IHRvIGJs b2NrIGRldmljZSB1bnRpbAo+IHRoZSBwYXRoIGlzIGZpeGVkPyBXb3VsZCBpdCBub3QgY2F1c2Ug ZmFpbHVyZXMgaWYgdGhlIGlzc3VlIHBlcnNpc3RzCj4gZm9yIHNlY29uZHM/IAoKQXMgSSBzYWlk LCB0aGF0J3Mgbm90IGhvdyBpdCB3b3Jrcy4gUjEtPlAxLCBSMi0+UDIsIC4uLiBvbmx5IGhvbGRz IGFzCmxvbmcgYXMgYWxsIHBhdGhzIGFyZSB1cCAoYW5kIG5vIElPIHNjaGVkdWxlciBpcyBhY3Rp dmUgd2hpY2ggbWlnaHQgcmUtCm9yZGVyIHlvdXIgSS9PIHJlcXVlc3RzKS4KCj4gV2hhdCBhYm91 dCB0aGUgc2l6ZSBvZiBxdWV1ZT8gSXMgdGhlcmUgYW55IGRhbmdlciBvZiBxdWV1ZSBnZXR0aW5n Cj4gb3ZlcmxvYWRlZD9BbnkgcG9pbnRlcnMgb3IgcmVmZXJlbmNlcyB3b3VsZCBiZSBvZiBncmVh dCBoZWxwLgoKVGhlb3JldGljYWxseSwgdGhlIHF1ZXVlIGlzIG9ubHkgbGltaXRlZCBieSBtZW1v cnkgc2l6ZS4KCk1hcnRpbgoKPiAKPiBUaGFua3MhCj4gS2FyYW4KPiAKPiAKPiBHZXQgT3V0bG9v ayBmb3IgQW5kcm9pZAo+IAo+IC0tCj4gZG0tZGV2ZWwgbWFpbGluZyBsaXN0Cj4gZG0tZGV2ZWxA cmVkaGF0LmNvbQo+IGh0dHBzOi8vd3d3LnJlZGhhdC5jb20vbWFpbG1hbi9saXN0aW5mby9kbS1k ZXZlbAoKLS0gCkRyLiBNYXJ0aW4gV2lsY2sgPG13aWxja0BzdXNlLmNvbT4sIFRlbC4gKzQ5ICgw KTkxMSA3NDA1MyAyMTA3ClNVU0UgTGludXggR21iSCwgR0Y6IEZlbGl4IEltZW5kw7ZyZmZlciwg SmFuZSBTbWl0aGFyZCwgR3JhaGFtIE5vcnRvbgpIUkIgMjEyODQgKEFHIE7DvHJuYmVyZykKCi0t CmRtLWRldmVsIG1haWxpbmcgbGlzdApkbS1kZXZlbEByZWRoYXQuY29tCmh0dHBzOi8vd3d3LnJl ZGhhdC5jb20vbWFpbG1hbi9saXN0aW5mby9kbS1kZXZlbA==