From mboxrd@z Thu Jan 1 00:00:00 1970 Content-Type: multipart/mixed; boundary="===============7785086606877588977==" MIME-Version: 1.0 From: Harris, James R Subject: [SPDK] Re: SPDK RAID5 support Date: Mon, 14 Oct 2019 17:43:27 +0000 Message-ID: In-Reply-To: 004701d581a8$4b613bd0$e223b370$@dev.mellanox.co.il List-ID: To: spdk@lists.01.org --===============7785086606877588977== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable =EF=BB=BFOn 10/13/19, 2:28 AM, "Sasha Kotchubievsky" wrote: Hi, = I'm very exiting to progress in RAID development in SPDK. = Artur, will you focus on RAID5 only, or RAID6 is also will be supported? = I think, it's important to keep existing SPDK approach for configuratio= n. = It would nice to see an abstraction for parity calculation. I believe, = the calculation can be optimized for specific platform, or even using HW ac= celerators is possible . [Jim] Agreed. To start we may use something similar to CRC calculations (= lib/util/crc32c.c), where we pick a parity calculation implementation at co= mpile time. Longer term, something more dynamic like the SPDK copy_engine = (for memcopies) or DPDK framework (for crypto/compression) would be even ni= cer. In case of distributed storage, even existing in market network cards c= an optimize RAID related operations. I believe, the next, upcoming generati= on will this capabilities to the next level. That needs some support from b= dev layer like events about configuration changes, or recovery/degradation. = = Best regards Sasha = -----Original Message----- From: =E6=9D=BE=E6=9C=AC=E5=91=A8=E5=B9=B3 / MATSUMOTO=EF=BC=8CSHUUHEI = = Sent: Friday, October 4, 2019 1:49 AM To: Storage Performance Development Kit Cc: Karkra, Kapil ; Baldysiak, Pawel ; Ptak, Slawomir Subject: [SPDK] Re: SPDK RAID5 support = Hi Artur, Paul, and All, = Thank you so much, I'm excited to know this. = Recently SPDK are starting to support DIF feature. Can we have any possibility to include the extended LBA (block size =3D= 512 + 8, 4096 + 128, or etc) into SPDK RAID? Do you have any comment? = Thanks, Shuhei = ________________________________ =E5=B7=AE=E5=87=BA=E4=BA=BA: Luse, Paul E =E9=80=81=E4=BF=A1=E6=97=A5=E6=99=82: 2019=E5=B9=B410=E6=9C=884=E6=97= =A5 4:20 =E5=AE=9B=E5=85=88: Storage Performance Development Kit CC: Karkra, Kapil ; Baldysiak, Pawel ; Ptak, Slawomir =E4=BB=B6=E5=90=8D: [SPDK] Re: SPDK RAID5 support = Hi Artur, = Thanks, I think this can be an awesome contribution. A few other thing= s to consider, you mention some of these already, so I just added some more= color. It would be good I think moving forward to put a trello board up w= ith a backlog of tasks so that others (like me __) can jump in and help als= o. I'd still like to get the RAID1E in there, but it makes little sense to= do it before any major refactoring. = I think it makes sense, if you guys are ready, to start putting togethe= r the backlog and knocking out a large series of small patches to get the r= efactoring done. = * the current RAID0 has no config on disk. We'll need to come up with a= scheme for handling existing RAID0 configured out of band with RAID5 using= COD. (like not allowing a RAID0 to be built on a set with COD, etc.) * the metadata layout, as you knee from working previous RAID projects,= needs to be thought out carefully to consider not only extensibility but v= ersion control for backwards compactivity and issues with conflicting COD t= hat are found. For example you have a 3 disk RAID5, take one of the disks = out and use it in another array somewhere then bring it back later and fire= up the original 3, the metadata has to have sufficient info to know who be= longs to what and which volumes to create and which to put in some sort of = offline state. I don't think we need that kind of capability right up front= (deciding how to deal with conflicts) but the metadata should have enough = information it up front. DDF was brought up before in the earlier RAID disc= ussions so just to make sure we're all on the same page, I see no value in = complying with that spec. Open to other thoughts though. * one thing to keep in mind, we probably want to retain common RPC code * we should keep migration in mind as well, not that it's something we = may ever need/want but lots of reserved space in metadata for tracking stat= e wrt migrations and rebuilds is needed. * similar to the RAID1E discussions earlier, there are some features yo= u mention as 'future' that most would consider a requirement for redundant = RAID - like degrade operation and rebuild. Those don't all have to go in at= the same time however without that minimum set we need to mark it experime= ntal, so nobody tries to use it thinking it has those basic things. * I can't remember if we talked about unit tests or not, but be sure th= e backlog includes getting solid UT coverage in there up front. The existi= ng UT code will likely need some refactoring as well to support the functio= n code refactoring. = -Paul = On 10/3/19, 3:00 AM, "Artur Paszkiewicz" wrote: = Hi all, = We want to add RAID5 support to SPDK. My team has experience with o= ther RAID projects, primarily with Linux MD RAID, which we actively develop a= nd support for Intel VROC. We already have an initial SPDK RAID5 implementatio= n created for an internal project. It has working read/write, including parti= al-stripe updates, parity calculation and reconstruct-reads. = Currently in SPDK there exists a RAID bdev module, which has only R= AID0 functionality. This can be used as a basis for a more generic RAID = stack. Here is our idea how to approach this: = 1. Refactor the bdev_raid module to separate RAID0-specific I/O han= dling code from more generic parts - configuration, bdev creation, etc. Move t= he RAID0 code to a new file. Use RAID level-specific callbacks, similar to e= xisting struct raid_fn_table. This architecture is also used in MD RAID dri= vers, where different RAID "personalities" work on top of a common layer. = 2. Add RAID5 support in another file, similarly to RAID0. Port our = current RAID5 code to this new framework. = 3. Incrementally add new functionalities. At this point, probably t= he most important will be support for member drive failure and degraded ope= ration, RAID rebuild and some form of on-disk metadata. = Any comments or suggestions are welcome. = Thanks, Artur _______________________________________________ SPDK mailing list -- spdk(a)lists.01.org To unsubscribe send an email to spdk-leave(a)lists.01.org = = _______________________________________________ SPDK mailing list -- spdk(a)lists.01.org To unsubscribe send an email to spdk-leave(a)lists.01.org _____________= __________________________________ SPDK mailing list -- spdk(a)lists.01.org To unsubscribe send an email to spdk-leave(a)lists.01.org _______________________________________________ SPDK mailing list -- spdk(a)lists.01.org To unsubscribe send an email to spdk-leave(a)lists.01.org = --===============7785086606877588977==--