Linux RAID subsystem development
 help / color / mirror / Atom feed
* [PATCH md 0 of 4] Introduction
@ 2004-08-23  3:10 NeilBrown
  2004-08-23  3:10 ` [PATCH md 1 of 4] Assorted fixes/improvemnet to generic md resync code NeilBrown
                   ` (3 more replies)
  0 siblings, 4 replies; 16+ messages in thread
From: NeilBrown @ 2004-08-23  3:10 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid


Following are 4 patches for md in 2.6.8.1-mm4

The first three are minor improvements and modifications either
required by or inspired by the fourth.

The fourth adds a new raid personality - raid10.  At 56K, I'm not 
sure it will get through the mailing list, but interested parties
can find it at:

  http://neilb.web.cse.unsw.edu.au/patches/linux-devel/2.6/2004-08-23-03

raid10 provides a combination of raid0 and raid1.
It requires mdadm 1.7.0 or later to use.  

The next release of mdadm should have better documention of raid10, but 
from the comment in the .c file:

/*
 * RAID10 provides a combination of RAID0 and RAID1 functionality.
 * The layout of data is defined by 
 *    chunk_size
 *    raid_disks
 *    near_copies (stored in low byte of layout)
 *    far_copies (stored in second byte of layout)
 *
 * The data to be stored is divided into chunks using chunksize.
 * Each device is divided into far_copies sections.
 * In each section, chunks are layed out in a style similar to raid0, but
 * near_copies copies of each chunk is stored (each on a different drive).
 * The starting device for each section is offset near_copies from the starting
 * device of the previous section.
 * Thus there are (near_copies*far_copies) of each chunk, and each is on a different
 * drive.
 * near_copies and far_copies must be at least one, and there product is at most
 * raid_disks.
 */

raid10 is currently marked EXPERIMENTAL, and this should be taken seriously.
A reasonable amount of basic testing hasn't shown any bugs, and it seems to resync
and rebuild correctly.  However wider testing would help.

NeilBrown

^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 2 of 4] Assorted minor md/raid1 fixes
  2004-08-23  3:10 [PATCH md 0 of 4] Introduction NeilBrown
                   ` (2 preceding siblings ...)
  2004-08-23  3:10 ` [PATCH md 3 of 4] Remove most calls to __bdevname from md.c NeilBrown
@ 2004-08-23  3:10 ` NeilBrown
  3 siblings, 0 replies; 16+ messages in thread
From: NeilBrown @ 2004-08-23  3:10 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid


1/ rationalise read_balance and "map" in raid1.  Discard map and 
   tidyup the interface to read_balance so it can be used instead.

2/ use offsetof rather than a caclulation to find the size of an
   structure with a var-length array at the end.

3/ remove some meaningless #defines 

4/ use printk_ratelimit to limit reports of failed sectors being redirected.

Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au>

### Diffstat output
 ./drivers/md/raid1.c |   89 ++++++++++++++++++++-------------------------------
 1 files changed, 36 insertions(+), 53 deletions(-)

diff ./drivers/md/raid1.c~current~ ./drivers/md/raid1.c
--- ./drivers/md/raid1.c~current~	2004-08-23 12:20:59.000000000 +1000
+++ ./drivers/md/raid1.c	2004-08-23 12:31:36.000000000 +1000
@@ -24,10 +24,6 @@
 
 #include <linux/raid/raid1.h>
 
-#define MAJOR_NR MD_MAJOR
-#define MD_DRIVER
-#define MD_PERSONALITY
-
 /*
  * Number of guaranteed r1bios in case of extreme VM load:
  */
@@ -44,13 +40,12 @@ static void * r1bio_pool_alloc(int gfp_f
 {
 	struct pool_info *pi = data;
 	r1bio_t *r1_bio;
+	int size = offsetof(r1bio_t, bios[pi->raid_disks]);
 
 	/* allocate a r1bio with room for raid_disks entries in the bios array */
-	r1_bio = kmalloc(sizeof(r1bio_t) + sizeof(struct bio*)*pi->raid_disks,
-			 gfp_flags);
+	r1_bio = kmalloc(size, gfp_flags);
 	if (r1_bio)
-		memset(r1_bio, 0, sizeof(*r1_bio) +
-			       sizeof(struct bio*) * pi->raid_disks);
+		memset(r1_bio, 0, size);
 	else
 		unplug_slaves(pi->mddev);
 
@@ -104,7 +99,7 @@ static void * r1buf_pool_alloc(int gfp_f
 		bio->bi_io_vec[i].bv_page = page;
 	}
 
-	r1_bio->master_bio = bio;
+	r1_bio->master_bio = NULL;
 
 	return r1_bio;
 
@@ -189,32 +184,6 @@ static inline void put_buf(r1bio_t *r1_b
 	spin_unlock_irqrestore(&conf->resync_lock, flags);
 }
 
-static int map(mddev_t *mddev, mdk_rdev_t **rdevp)
-{
-	conf_t *conf = mddev_to_conf(mddev);
-	int i, disks = conf->raid_disks;
-
-	/*
-	 * Later we do read balancing on the read side
-	 * now we use the first available disk.
-	 */
-
-	spin_lock_irq(&conf->device_lock);
-	for (i = 0; i < disks; i++) {
-		mdk_rdev_t *rdev = conf->mirrors[i].rdev;
-		if (rdev && rdev->in_sync) {
-			*rdevp = rdev;
-			atomic_inc(&rdev->nr_pending);
-			spin_unlock_irq(&conf->device_lock);
-			return i;
-		}
-	}
-	spin_unlock_irq(&conf->device_lock);
-
-	printk(KERN_ERR "raid1_map(): huh, no more operational devices?\n");
-	return -1;
-}
-
 static void reschedule_retry(r1bio_t *r1_bio)
 {
 	unsigned long flags;
@@ -292,8 +261,9 @@ static int raid1_end_read_request(struct
 		 * oops, read error:
 		 */
 		char b[BDEVNAME_SIZE];
-		printk(KERN_ERR "raid1: %s: rescheduling sector %llu\n",
-		       bdevname(conf->mirrors[mirror].rdev->bdev,b), (unsigned long long)r1_bio->sector);
+		if (printk_ratelimit())
+			printk(KERN_ERR "raid1: %s: rescheduling sector %llu\n",
+			       bdevname(conf->mirrors[mirror].rdev->bdev,b), (unsigned long long)r1_bio->sector);
 		reschedule_retry(r1_bio);
 	}
 
@@ -363,11 +333,11 @@ static int raid1_end_write_request(struc
  *
  * The rdev for the device selected will have nr_pending incremented.
  */
-static int read_balance(conf_t *conf, struct bio *bio, r1bio_t *r1_bio)
+static int read_balance(conf_t *conf, r1bio_t *r1_bio)
 {
 	const unsigned long this_sector = r1_bio->sector;
 	int new_disk = conf->last_used, disk = new_disk;
-	const int sectors = bio->bi_size >> 9;
+	const int sectors = r1_bio->sectors;
 	sector_t new_distance, current_distance;
 
 	spin_lock_irq(&conf->device_lock);
@@ -378,14 +348,14 @@ static int read_balance(conf_t *conf, st
 	 */
 	if (conf->mddev->recovery_cp < MaxSector &&
 	    (this_sector + sectors >= conf->next_resync)) {
-		/* make sure that disk is operational */
+		/* Choose the first operation device, for consistancy */
 		new_disk = 0;
 
 		while (!conf->mirrors[new_disk].rdev ||
 		       !conf->mirrors[new_disk].rdev->in_sync) {
 			new_disk++;
 			if (new_disk == conf->raid_disks) {
-				new_disk = 0;
+				new_disk = -1;
 				break;
 			}
 		}
@@ -400,7 +370,7 @@ static int read_balance(conf_t *conf, st
 			new_disk = conf->raid_disks;
 		new_disk--;
 		if (new_disk == disk) {
-			new_disk = conf->last_used;
+			new_disk = -1;
 			goto rb_out;
 		}
 	}
@@ -440,13 +410,13 @@ static int read_balance(conf_t *conf, st
 	} while (disk != conf->last_used);
 
 rb_out:
-	r1_bio->read_disk = new_disk;
-	conf->next_seq_sect = this_sector + sectors;
 
-	conf->last_used = new_disk;
 
-	if (conf->mirrors[new_disk].rdev)
+	if (new_disk >= 0) {
+		conf->next_seq_sect = this_sector + sectors;
+		conf->last_used = new_disk;
 		atomic_inc(&conf->mirrors[new_disk].rdev->nr_pending);
+	}
 	spin_unlock_irq(&conf->device_lock);
 
 	return new_disk;
@@ -571,15 +541,26 @@ static int make_request(request_queue_t 
 	r1_bio->mddev = mddev;
 	r1_bio->sector = bio->bi_sector;
 
+	r1_bio->state = 0;
+
 	if (bio_data_dir(bio) == READ) {
 		/*
 		 * read balancing logic:
 		 */
-		mirror = conf->mirrors + read_balance(conf, bio, r1_bio);
+		int rdisk = read_balance(conf, r1_bio);
+
+		if (rdisk < 0) {
+			/* couldn't find anywhere to read from */
+			raid_end_bio_io(r1_bio);
+			return 0;
+		}
+		mirror = conf->mirrors + rdisk;
+
+		r1_bio->read_disk = rdisk;
 
 		read_bio = bio_clone(bio, GFP_NOIO);
 
-		r1_bio->bios[r1_bio->read_disk] = read_bio;
+		r1_bio->bios[rdisk] = read_bio;
 
 		read_bio->bi_sector = r1_bio->sector + mirror->rdev->data_offset;
 		read_bio->bi_bdev = mirror->rdev->bdev;
@@ -951,7 +932,7 @@ static void raid1d(mddev_t *mddev)
 		} else {
 			int disk;
 			bio = r1_bio->bios[r1_bio->read_disk];
-			if ((disk=map(mddev, &rdev)) == -1) {
+			if ((disk=read_balance(conf, r1_bio)) == -1) {
 				printk(KERN_ALERT "raid1: %s: unrecoverable I/O"
 				       " read error for block %llu\n",
 				       bdevname(bio->bi_bdev,b),
@@ -961,10 +942,12 @@ static void raid1d(mddev_t *mddev)
 				r1_bio->bios[r1_bio->read_disk] = NULL;
 				r1_bio->read_disk = disk;
 				r1_bio->bios[r1_bio->read_disk] = bio;
-				printk(KERN_ERR "raid1: %s: redirecting sector %llu to"
-				       " another mirror\n",
-				       bdevname(rdev->bdev,b),
-				       (unsigned long long)r1_bio->sector);
+				rdev = conf->mirrors[disk].rdev;
+				if (printk_ratelimit())
+					printk(KERN_ERR "raid1: %s: redirecting sector %llu to"
+					       " another mirror\n",
+					       bdevname(rdev->bdev,b),
+					       (unsigned long long)r1_bio->sector);
 				bio->bi_bdev = rdev->bdev;
 				bio->bi_sector = r1_bio->sector + rdev->data_offset;
 				bio->bi_rw = READ;

^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 4 of 4] RAID10 module for MD
  2004-08-23  3:10 [PATCH md 0 of 4] Introduction NeilBrown
  2004-08-23  3:10 ` [PATCH md 1 of 4] Assorted fixes/improvemnet to generic md resync code NeilBrown
@ 2004-08-23  3:10 ` NeilBrown
  2004-08-23  3:10 ` [PATCH md 3 of 4] Remove most calls to __bdevname from md.c NeilBrown
  2004-08-23  3:10 ` [PATCH md 2 of 4] Assorted minor md/raid1 fixes NeilBrown
  3 siblings, 0 replies; 16+ messages in thread
From: NeilBrown @ 2004-08-23  3:10 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid


This patch adds a 'raid10' module which provides
features similar to both raid0 and raid1 in the one
array.  Various combinations of layout are supported.

This code is still "experimental", but appears to work.
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au>

### Diffstat output
 ./drivers/md/Kconfig          |   18 
 ./drivers/md/Makefile         |    1 
 ./drivers/md/raid10.c         | 1780 ++++++++++++++++++++++++++++++++++++++++++
 ./include/linux/raid/md_k.h   |    5 
 ./include/linux/raid/raid10.h |  103 ++
 5 files changed, 1906 insertions(+), 1 deletion(-)

diff ./drivers/md/Kconfig~current~ ./drivers/md/Kconfig
--- ./drivers/md/Kconfig~current~	2004-08-23 13:06:02.000000000 +1000
+++ ./drivers/md/Kconfig	2004-08-23 12:49:27.000000000 +1000
@@ -85,6 +85,24 @@ config MD_RAID1
 
 	  If unsure, say Y.
 
+config MD_RAID10
+	tristate "RAID-10 (mirrored striping) mode (EXPERIMENTAL)"
+	depends on BLK_DEV_MD && EXPERIMENTAL
+	---help---
+	  RAID-10 provides a combination of striping (RAID-0) and
+	  mirroring (RAID-1) with easier configuration and more flexable
+	  layout.
+	  Unlike RAID-0, but like RAID-1, RAID-10 requires all devices to
+	  be the same size (or atleast, only as much as the smallest device
+	  will be used).
+	  RAID-10 provides a variety of layouts that provide different levels
+	  of redundancy and performance.
+	
+	  RAID-10 requires mdadm-1.7.0 or later, available at:
+
+	  ftp://ftp.kernel.org/pub/linux/utils/raid/mdadm/
+
+
 config MD_RAID5
 	tristate "RAID-4/RAID-5 mode"
 	depends on BLK_DEV_MD

diff ./drivers/md/Makefile~current~ ./drivers/md/Makefile
--- ./drivers/md/Makefile~current~	2004-08-23 13:06:02.000000000 +1000
+++ ./drivers/md/Makefile	2004-08-23 12:49:27.000000000 +1000
@@ -20,6 +20,7 @@ hostprogs-y	:= mktables
 obj-$(CONFIG_MD_LINEAR)		+= linear.o
 obj-$(CONFIG_MD_RAID0)		+= raid0.o
 obj-$(CONFIG_MD_RAID1)		+= raid1.o
+obj-$(CONFIG_MD_RAID10)		+= raid10.o
 obj-$(CONFIG_MD_RAID5)		+= raid5.o xor.o
 obj-$(CONFIG_MD_RAID6)		+= raid6.o xor.o
 obj-$(CONFIG_MD_MULTIPATH)	+= multipath.o

diff ./drivers/md/raid10.c~current~ ./drivers/md/raid10.c
--- ./drivers/md/raid10.c~current~	2004-08-23 13:06:02.000000000 +1000
+++ ./drivers/md/raid10.c	2004-08-23 13:06:30.000000000 +1000
@@ -0,0 +1,1780 @@
+/*
+ * raid10.c : Multiple Devices driver for Linux
+ *
+ * Copyright (C) 2000-2004 Neil Brown
+ *
+ * RAID-10 support for md.
+ *
+ * Base on code in raid1.c.  See raid1.c for futher copyright information.
+ *
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2, or (at your option)
+ * any later version.
+ *
+ * You should have received a copy of the GNU General Public License
+ * (for example /usr/src/linux/COPYING); if not, write to the Free
+ * Software Foundation, Inc., 675 Mass Ave, Cambridge, MA 02139, USA.
+ */
+
+#include <linux/raid/raid10.h>
+
+/*
+ * RAID10 provides a combination of RAID0 and RAID1 functionality.
+ * The layout of data is defined by 
+ *    chunk_size
+ *    raid_disks
+ *    near_copies (stored in low byte of layout)
+ *    far_copies (stored in second byte of layout)
+ *
+ * The data to be stored is divided into chunks using chunksize.
+ * Each device is divided into far_copies sections.
+ * In each section, chunks are laid out in a style similar to raid0, but
+ * near_copies copies of each chunk is stored (each on a different drive).
+ * The starting device for each section is offset near_copies from the starting
+ * device of the previous section.
+ * Thus there are (near_copies*far_copies) of each chunk, and each is on a different
+ * drive.
+ * near_copies and far_copies must be at least one, and there product is at most
+ * raid_disks.
+ */
+
+/*
+ * Number of guaranteed r10bios in case of extreme VM load:
+ */
+#define	NR_RAID10_BIOS 256
+
+static void unplug_slaves(mddev_t *mddev);
+
+static void * r10bio_pool_alloc(int gfp_flags, void *data)
+{
+	conf_t *conf = data;
+	r10bio_t *r10_bio;
+	int size = offsetof(struct r10bio_s, devs[conf->copies]);
+
+	/* allocate a r10bio with room for raid_disks entries in the bios array */
+	r10_bio = kmalloc(size, gfp_flags);
+	if (r10_bio)
+		memset(r10_bio, 0, size);
+	else
+		unplug_slaves(conf->mddev);
+
+	return r10_bio;
+}
+
+static void r10bio_pool_free(void *r10_bio, void *data)
+{
+	kfree(r10_bio);
+}
+
+#define RESYNC_BLOCK_SIZE (64*1024)
+//#define RESYNC_BLOCK_SIZE PAGE_SIZE
+#define RESYNC_SECTORS (RESYNC_BLOCK_SIZE >> 9)
+#define RESYNC_PAGES ((RESYNC_BLOCK_SIZE + PAGE_SIZE-1) / PAGE_SIZE)
+#define RESYNC_WINDOW (2048*1024)
+
+/*
+ * When performing a resync, we need to read and compare, so
+ * we need as many pages are there are copies.
+ * When performing a recovery, we need 2 bios, one for read,
+ * one for write (we recover only one drive per r10buf)
+ *
+ */
+static void * r10buf_pool_alloc(int gfp_flags, void *data)
+{
+	conf_t *conf = data;
+	struct page *page;
+	r10bio_t *r10_bio;
+	struct bio *bio;
+	int i, j;
+	int nalloc;
+
+	r10_bio = r10bio_pool_alloc(gfp_flags, conf);
+	if (!r10_bio) {
+		unplug_slaves(conf->mddev);
+		return NULL;
+	}
+
+	if (test_bit(MD_RECOVERY_SYNC, &conf->mddev->recovery))
+		nalloc = conf->copies; /* resync */
+	else
+		nalloc = 2; /* recovery */
+
+	/*
+	 * Allocate bios.
+	 */
+	for (j = nalloc ; j-- ; ) {
+		bio = bio_alloc(gfp_flags, RESYNC_PAGES);
+		if (!bio)
+			goto out_free_bio;
+		r10_bio->devs[j].bio = bio;
+	}
+	/*
+	 * Allocate RESYNC_PAGES data pages and attach them
+	 * where needed.
+	 */
+	for (j = 0 ; j < nalloc; j++) {
+		bio = r10_bio->devs[j].bio;
+		for (i = 0; i < RESYNC_PAGES; i++) {
+			page = alloc_page(gfp_flags);
+			if (unlikely(!page))
+				goto out_free_pages;
+
+			bio->bi_io_vec[i].bv_page = page;
+		}
+	}
+
+	return r10_bio;
+
+out_free_pages:
+	for ( ; i > 0 ; i--)
+		__free_page(bio->bi_io_vec[i-1].bv_page);
+	while (j--)
+		for (i = 0; i < RESYNC_PAGES ; i++)
+			__free_page(r10_bio->devs[j].bio->bi_io_vec[i].bv_page);
+	j = -1;
+out_free_bio:
+	while ( ++j < nalloc )
+		bio_put(r10_bio->devs[j].bio);
+	r10bio_pool_free(r10_bio, conf);
+	return NULL;
+}
+
+static void r10buf_pool_free(void *__r10_bio, void *data)
+{
+	int i;
+	conf_t *conf = data;
+	r10bio_t *r10bio = __r10_bio;
+	int j;
+
+	for (j=0; j < conf->copies; j++) {
+		struct bio *bio = r10bio->devs[j].bio;
+		if (bio) {
+			for (i = 0; i < RESYNC_PAGES; i++) {
+				__free_page(bio->bi_io_vec[i].bv_page);
+				bio->bi_io_vec[i].bv_page = NULL;
+			}
+			bio_put(bio);
+		}
+	}
+	r10bio_pool_free(r10bio, conf);
+}
+
+static void put_all_bios(conf_t *conf, r10bio_t *r10_bio)
+{
+	int i;
+
+	for (i = 0; i < conf->copies; i++) {
+		struct bio **bio = & r10_bio->devs[i].bio;
+		if (*bio)
+			bio_put(*bio);
+		*bio = NULL;
+	}
+}
+
+static inline void free_r10bio(r10bio_t *r10_bio)
+{
+	unsigned long flags;
+
+	conf_t *conf = mddev_to_conf(r10_bio->mddev);
+
+	/*
+	 * Wake up any possible resync thread that waits for the device
+	 * to go idle.
+	 */
+	spin_lock_irqsave(&conf->resync_lock, flags);
+	if (!--conf->nr_pending) {
+		wake_up(&conf->wait_idle);
+		wake_up(&conf->wait_resume);
+	}
+	spin_unlock_irqrestore(&conf->resync_lock, flags);
+
+	put_all_bios(conf, r10_bio);
+	mempool_free(r10_bio, conf->r10bio_pool);
+}
+
+static inline void put_buf(r10bio_t *r10_bio)
+{
+	conf_t *conf = mddev_to_conf(r10_bio->mddev);
+	unsigned long flags;
+
+	mempool_free(r10_bio, conf->r10buf_pool);
+
+	spin_lock_irqsave(&conf->resync_lock, flags);
+	if (!conf->barrier)
+		BUG();
+	--conf->barrier;
+	wake_up(&conf->wait_resume);
+	wake_up(&conf->wait_idle);
+
+	if (!--conf->nr_pending) {
+		wake_up(&conf->wait_idle);
+		wake_up(&conf->wait_resume);
+	}
+	spin_unlock_irqrestore(&conf->resync_lock, flags);
+}
+
+static void reschedule_retry(r10bio_t *r10_bio)
+{
+	unsigned long flags;
+	mddev_t *mddev = r10_bio->mddev;
+	conf_t *conf = mddev_to_conf(mddev);
+
+	spin_lock_irqsave(&conf->device_lock, flags);
+	list_add(&r10_bio->retry_list, &conf->retry_list);
+	spin_unlock_irqrestore(&conf->device_lock, flags);
+
+	md_wakeup_thread(mddev->thread);
+}
+
+/*
+ * raid_end_bio_io() is called when we have finished servicing a mirrored
+ * operation and are ready to return a success/failure code to the buffer
+ * cache layer.
+ */
+static void raid_end_bio_io(r10bio_t *r10_bio)
+{
+	struct bio *bio = r10_bio->master_bio;
+
+	bio_endio(bio, bio->bi_size,
+		test_bit(R10BIO_Uptodate, &r10_bio->state) ? 0 : -EIO);
+	free_r10bio(r10_bio);
+}
+
+/*
+ * Update disk head position estimator based on IRQ completion info.
+ */
+static inline void update_head_pos(int slot, r10bio_t *r10_bio)
+{
+	conf_t *conf = mddev_to_conf(r10_bio->mddev);
+
+	conf->mirrors[r10_bio->devs[slot].devnum].head_position =
+		r10_bio->devs[slot].addr + (r10_bio->sectors);
+}
+
+static int raid10_end_read_request(struct bio *bio, unsigned int bytes_done, int error)
+{
+	int uptodate = test_bit(BIO_UPTODATE, &bio->bi_flags);
+	r10bio_t * r10_bio = (r10bio_t *)(bio->bi_private);
+	int slot, dev;
+	conf_t *conf = mddev_to_conf(r10_bio->mddev);
+
+	if (bio->bi_size)
+		return 1;
+
+	slot = r10_bio->read_slot;
+	dev = r10_bio->devs[slot].devnum;
+	/*
+	 * this branch is our 'one mirror IO has finished' event handler:
+	 */
+	if (!uptodate)
+		md_error(r10_bio->mddev, conf->mirrors[dev].rdev);
+	else
+		/*
+		 * Set R10BIO_Uptodate in our master bio, so that
+		 * we will return a good error code to the higher
+		 * levels even if IO on some other mirrored buffer fails.
+		 *
+		 * The 'master' represents the composite IO operation to
+		 * user-side. So if something waits for IO, then it will
+		 * wait for the 'master' bio.
+		 */
+		set_bit(R10BIO_Uptodate, &r10_bio->state);
+
+	update_head_pos(slot, r10_bio);
+
+	/*
+	 * we have only one bio on the read side
+	 */
+	if (uptodate)
+		raid_end_bio_io(r10_bio);
+	else {
+		/*
+		 * oops, read error:
+		 */
+		char b[BDEVNAME_SIZE];
+		if (printk_ratelimit())
+			printk(KERN_ERR "raid10: %s: rescheduling sector %llu\n",
+			       bdevname(conf->mirrors[dev].rdev->bdev,b), (unsigned long long)r10_bio->sector);
+		reschedule_retry(r10_bio);
+	}
+
+	rdev_dec_pending(conf->mirrors[dev].rdev, conf->mddev);
+	return 0;
+}
+
+static int raid10_end_write_request(struct bio *bio, unsigned int bytes_done, int error)
+{
+	int uptodate = test_bit(BIO_UPTODATE, &bio->bi_flags);
+	r10bio_t * r10_bio = (r10bio_t *)(bio->bi_private);
+	int slot, dev;
+	conf_t *conf = mddev_to_conf(r10_bio->mddev);
+
+	if (bio->bi_size)
+		return 1;
+
+	for (slot = 0; slot < conf->copies; slot++)
+		if (r10_bio->devs[slot].bio == bio)
+			break;
+	dev = r10_bio->devs[slot].devnum;
+
+	/*
+	 * this branch is our 'one mirror IO has finished' event handler:
+	 */
+	if (!uptodate)
+		md_error(r10_bio->mddev, conf->mirrors[dev].rdev);
+	else
+		/*
+		 * Set R10BIO_Uptodate in our master bio, so that
+		 * we will return a good error code for to the higher
+		 * levels even if IO on some other mirrored buffer fails.
+		 *
+		 * The 'master' represents the composite IO operation to
+		 * user-side. So if something waits for IO, then it will
+		 * wait for the 'master' bio.
+		 */
+		set_bit(R10BIO_Uptodate, &r10_bio->state);
+
+	update_head_pos(slot, r10_bio);
+
+	/*
+	 *
+	 * Let's see if all mirrored write operations have finished
+	 * already.
+	 */
+	if (atomic_dec_and_test(&r10_bio->remaining)) {
+		md_write_end(r10_bio->mddev);
+		raid_end_bio_io(r10_bio);
+	}
+
+	rdev_dec_pending(conf->mirrors[dev].rdev, conf->mddev);
+	return 0;
+}
+
+
+/*
+ * RAID10 layout manager
+ * Aswell as the chunksize and raid_disks count, there are two
+ * parameters: near_copies and far_copies.
+ * near_copies * far_copies must be <= raid_disks.
+ * Normally one of these will be 1.
+ * If both are 1, we get raid0.
+ * If near_copies == raid_disks, we get raid1.
+ *
+ * Chunks are layed out in raid0 style with near_copies copies of the
+ * first chunk, followed by near_copies copies of the next chunk and
+ * so on.
+ * If far_copies > 1, then after 1/far_copies of the array has been assigned
+ * as described above, we start again with a device offset of near_copies.
+ * So we effectively have another copy of the whole array further down all
+ * the drives, but with blocks on different drives.
+ * With this layout, and block is never stored twice on the one device.
+ *
+ * raid10_find_phys finds the sector offset of a given virtual sector
+ * on each device that it is on. If a block isn't on a device,
+ * that entry in the array is set to MaxSector.
+ *
+ * raid10_find_virt does the reverse mapping, from a device and a
+ * sector offset to a virtual address
+ */
+
+static void raid10_find_phys(conf_t *conf, r10bio_t *r10bio)
+{
+	int n,f;
+	sector_t sector;
+	sector_t chunk;
+	sector_t stripe;
+	int dev;
+
+	int slot = 0;
+
+	/* now calculate first sector/dev */
+	chunk = r10bio->sector >> conf->chunk_shift;
+	sector = r10bio->sector & conf->chunk_mask;
+
+	chunk *= conf->near_copies;
+	stripe = chunk;
+	dev = sector_div(stripe, conf->raid_disks);
+
+	sector += stripe << conf->chunk_shift;
+
+	/* and calculate all the others */
+	for (n=0; n < conf->near_copies; n++) {
+		int d = dev;
+		sector_t s = sector;
+		r10bio->devs[slot].addr = sector;
+		r10bio->devs[slot].devnum = d;
+		slot++;
+
+		for (f = 1; f < conf->far_copies; f++) {
+			d += conf->near_copies;
+			if (d >= conf->raid_disks)
+				d -= conf->raid_disks;
+			s += conf->stride;
+			r10bio->devs[slot].devnum = d;
+			r10bio->devs[slot].addr = s;
+			slot++;
+		}
+		dev++;
+		if (dev >= conf->raid_disks) {
+			dev = 0;
+			sector += (conf->chunk_mask + 1);
+		}
+	}
+	BUG_ON(slot != conf->copies);
+}
+
+static sector_t raid10_find_virt(conf_t *conf, sector_t sector, int dev)
+{
+	sector_t offset, chunk, vchunk;
+
+	while (sector > conf->stride) {
+		sector -= conf->stride;
+		if (dev < conf->near_copies)
+			dev += conf->raid_disks - conf->near_copies;
+		else
+			dev -= conf->near_copies;
+	}
+
+	offset = sector & conf->chunk_mask;
+	chunk = sector >> conf->chunk_shift;
+	vchunk = chunk * conf->raid_disks + dev;
+	sector_div(vchunk, conf->near_copies);
+	return (vchunk << conf->chunk_shift) + offset;
+}
+
+/**
+ *	raid10_mergeable_bvec -- tell bio layer if a two requests can be merged
+ *	@q: request queue
+ *	@bio: the buffer head that's been built up so far
+ *	@biovec: the request that could be merged to it.
+ *
+ *	Return amount of bytes we can accept at this offset
+ *      If near_copies == raid_disk, there are no striping issues,
+ *      but in that case, the function isn't called at all.
+ */
+static int raid10_mergeable_bvec(request_queue_t *q, struct bio *bio,
+				struct bio_vec *bio_vec)
+{
+	mddev_t *mddev = q->queuedata;
+	sector_t sector = bio->bi_sector + get_start_sect(bio->bi_bdev);
+	int max;
+	unsigned int chunk_sectors = mddev->chunk_size >> 9;
+	unsigned int bio_sectors = bio->bi_size >> 9;
+
+	max =  (chunk_sectors - ((sector & (chunk_sectors - 1)) + bio_sectors)) << 9;
+	if (max < 0) max = 0; /* bio_add cannot handle a negative return */
+	if (max <= bio_vec->bv_len && bio_sectors == 0)
+		return bio_vec->bv_len;
+	else
+		return max;
+}
+
+/*
+ * This routine returns the disk from which the requested read should
+ * be done. There is a per-array 'next expected sequential IO' sector
+ * number - if this matches on the next IO then we use the last disk.
+ * There is also a per-disk 'last know head position' sector that is
+ * maintained from IRQ contexts, both the normal and the resync IO
+ * completion handlers update this position correctly. If there is no
+ * perfect sequential match then we pick the disk whose head is closest.
+ *
+ * If there are 2 mirrors in the same 2 devices, performance degrades
+ * because position is mirror, not device based.
+ *
+ * The rdev for the device selected will have nr_pending incremented.
+ */
+
+/*
+ * FIXME: possibly should rethink readbalancing and do it differently
+ * depending on near_copies / far_copies geometry.
+ */
+static int read_balance(conf_t *conf, r10bio_t *r10_bio)
+{
+	const unsigned long this_sector = r10_bio->sector;
+	int disk, slot, nslot;
+	const int sectors = r10_bio->sectors;
+	sector_t new_distance, current_distance;
+
+	raid10_find_phys(conf, r10_bio);
+	spin_lock_irq(&conf->device_lock);
+	/*
+	 * Check if we can balance. We can balance on the whole
+	 * device if no resync is going on, or below the resync window.
+	 * We take the first readable disk when above the resync window.
+	 */
+	if (conf->mddev->recovery_cp < MaxSector
+	    && (this_sector + sectors >= conf->next_resync)) {
+		/* make sure that disk is operational */
+		slot = 0;
+		disk = r10_bio->devs[slot].devnum;
+
+		while (!conf->mirrors[disk].rdev ||
+		       !conf->mirrors[disk].rdev->in_sync) {
+			slot++;
+			if (slot == conf->copies) {
+				slot = 0;
+				disk = -1;
+				break;
+			}
+			disk = r10_bio->devs[slot].devnum;
+		}
+		goto rb_out;
+	}
+
+
+	/* make sure the disk is operational */
+	slot = 0;
+	disk = r10_bio->devs[slot].devnum;
+	while (!conf->mirrors[disk].rdev ||
+	       !conf->mirrors[disk].rdev->in_sync) {
+		slot ++;
+		if (slot == conf->copies) {
+			disk = -1;
+			goto rb_out;
+		}
+		disk = r10_bio->devs[slot].devnum;
+	}
+
+
+	current_distance = abs(this_sector - conf->mirrors[disk].head_position);
+
+	/* Find the disk whose head is closest */
+
+	for (nslot = slot; nslot < conf->copies; nslot++) {
+		int ndisk = r10_bio->devs[nslot].devnum;
+
+
+		if (!conf->mirrors[ndisk].rdev ||
+		    !conf->mirrors[ndisk].rdev->in_sync)
+			continue;
+
+		if (!atomic_read(&conf->mirrors[ndisk].rdev->nr_pending)) {
+			disk = ndisk;
+			slot = nslot;
+			break;
+		}
+		new_distance = abs(r10_bio->devs[nslot].addr -
+				   conf->mirrors[ndisk].head_position);
+		if (new_distance < current_distance) {
+			current_distance = new_distance;
+			disk = ndisk;
+			slot = nslot;
+		}
+	}
+
+rb_out:
+	r10_bio->read_slot = slot;
+/*	conf->next_seq_sect = this_sector + sectors;*/
+
+	if (disk >= 0 && conf->mirrors[disk].rdev)
+		atomic_inc(&conf->mirrors[disk].rdev->nr_pending);
+	spin_unlock_irq(&conf->device_lock);
+
+	return disk;
+}
+
+static void unplug_slaves(mddev_t *mddev)
+{
+	conf_t *conf = mddev_to_conf(mddev);
+	int i;
+	unsigned long flags;
+
+	spin_lock_irqsave(&conf->device_lock, flags);
+	for (i=0; i<mddev->raid_disks; i++) {
+		mdk_rdev_t *rdev = conf->mirrors[i].rdev;
+		if (rdev && atomic_read(&rdev->nr_pending)) {
+			request_queue_t *r_queue = bdev_get_queue(rdev->bdev);
+
+			atomic_inc(&rdev->nr_pending);
+			spin_unlock_irqrestore(&conf->device_lock, flags);
+
+			if (r_queue->unplug_fn)
+				r_queue->unplug_fn(r_queue);
+
+			spin_lock_irqsave(&conf->device_lock, flags);
+			atomic_dec(&rdev->nr_pending);
+		}
+	}
+	spin_unlock_irqrestore(&conf->device_lock, flags);
+}
+static void raid10_unplug(request_queue_t *q)
+{
+	unplug_slaves(q->queuedata);
+}
+
+static int raid10_issue_flush(request_queue_t *q, struct gendisk *disk,
+			     sector_t *error_sector)
+{
+	mddev_t *mddev = q->queuedata;
+	conf_t *conf = mddev_to_conf(mddev);
+	unsigned long flags;
+	int i, ret = 0;
+
+	spin_lock_irqsave(&conf->device_lock, flags);
+	for (i=0; i<mddev->raid_disks; i++) {
+		mdk_rdev_t *rdev = conf->mirrors[i].rdev;
+		if (rdev && !rdev->faulty) {
+			struct block_device *bdev = rdev->bdev;
+			request_queue_t *r_queue = bdev_get_queue(bdev);
+
+			if (r_queue->issue_flush_fn) {
+				ret = r_queue->issue_flush_fn(r_queue, bdev->bd_disk, error_sector);
+				if (ret)
+					break;
+			}
+		}
+	}
+	spin_unlock_irqrestore(&conf->device_lock, flags);
+	return ret;
+}
+
+/*
+ * Throttle resync depth, so that we can both get proper overlapping of
+ * requests, but are still able to handle normal requests quickly.
+ */
+#define RESYNC_DEPTH 32
+
+static void device_barrier(conf_t *conf, sector_t sect)
+{
+	spin_lock_irq(&conf->resync_lock);
+	wait_event_lock_irq(conf->wait_idle, !waitqueue_active(&conf->wait_resume),
+			    conf->resync_lock, unplug_slaves(conf->mddev));
+
+	if (!conf->barrier++) {
+		wait_event_lock_irq(conf->wait_idle, !conf->nr_pending,
+				    conf->resync_lock, unplug_slaves(conf->mddev));
+		if (conf->nr_pending)
+			BUG();
+	}
+	wait_event_lock_irq(conf->wait_resume, conf->barrier < RESYNC_DEPTH,
+			    conf->resync_lock, unplug_slaves(conf->mddev));
+	conf->next_resync = sect;
+	spin_unlock_irq(&conf->resync_lock);
+}
+
+static int make_request(request_queue_t *q, struct bio * bio)
+{
+	mddev_t *mddev = q->queuedata;
+	conf_t *conf = mddev_to_conf(mddev);
+	mirror_info_t *mirror;
+	r10bio_t *r10_bio;
+	struct bio *read_bio;
+	int i;
+	int chunk_sects = conf->chunk_mask + 1;
+
+	/* If this request crosses a chunk boundary, we need to
+	 * split it.  This will only happen for 1 PAGE (or less) requests.
+	 */
+	if (unlikely( (bio->bi_sector & conf->chunk_mask) + (bio->bi_size >> 9)
+		      > chunk_sects &&
+		    conf->near_copies < conf->raid_disks)) {
+		struct bio_pair *bp;
+		/* Sanity check -- queue functions should prevent this happening */
+		if (bio->bi_vcnt != 1 ||
+		    bio->bi_idx != 0)
+			goto bad_map;
+		/* This is a one page bio that upper layers
+		 * refuse to split for us, so we need to split it.
+		 */
+		bp = bio_split(bio, bio_split_pool,
+			       chunk_sects - (bio->bi_sector & (chunk_sects - 1)) );
+		if (make_request(q, &bp->bio1))
+			generic_make_request(&bp->bio1);
+		if (make_request(q, &bp->bio2))
+			generic_make_request(&bp->bio2);
+
+		bio_pair_release(bp);
+		return 0;
+	bad_map:
+		printk("raid10_make_request bug: can't convert block across chunks"
+		       " or bigger than %dk %llu %d\n", chunk_sects/2,
+		       (unsigned long long)bio->bi_sector, bio->bi_size >> 10);
+
+		bio_io_error(bio, bio->bi_size);
+		return 0;
+	}
+
+	/*
+	 * Register the new request and wait if the reconstruction
+	 * thread has put up a bar for new requests.
+	 * Continue immediately if no resync is active currently.
+	 */
+	spin_lock_irq(&conf->resync_lock);
+	wait_event_lock_irq(conf->wait_resume, !conf->barrier, conf->resync_lock, );
+	conf->nr_pending++;
+	spin_unlock_irq(&conf->resync_lock);
+
+	if (bio_data_dir(bio)==WRITE) {
+		disk_stat_inc(mddev->gendisk, writes);
+		disk_stat_add(mddev->gendisk, write_sectors, bio_sectors(bio));
+	} else {
+		disk_stat_inc(mddev->gendisk, reads);
+		disk_stat_add(mddev->gendisk, read_sectors, bio_sectors(bio));
+	}
+
+	r10_bio = mempool_alloc(conf->r10bio_pool, GFP_NOIO);
+
+	r10_bio->master_bio = bio;
+	r10_bio->sectors = bio->bi_size >> 9;
+
+	r10_bio->mddev = mddev;
+	r10_bio->sector = bio->bi_sector;
+
+	if (bio_data_dir(bio) == READ) {
+		/*
+		 * read balancing logic:
+		 */
+		int disk = read_balance(conf, r10_bio);
+		int slot = r10_bio->read_slot;
+		if (disk < 0) {
+			raid_end_bio_io(r10_bio);
+			return 0;
+		}
+		mirror = conf->mirrors + disk;
+
+		read_bio = bio_clone(bio, GFP_NOIO);
+
+		r10_bio->devs[slot].bio = read_bio;
+
+		read_bio->bi_sector = r10_bio->devs[slot].addr +
+			mirror->rdev->data_offset;
+		read_bio->bi_bdev = mirror->rdev->bdev;
+		read_bio->bi_end_io = raid10_end_read_request;
+		read_bio->bi_rw = READ;
+		read_bio->bi_private = r10_bio;
+
+		generic_make_request(read_bio);
+		return 0;
+	}
+
+	/*
+	 * WRITE:
+	 */
+	/* first select target devices under spinlock and
+	 * inc refcount on their rdev.  Record them by setting
+	 * bios[x] to bio
+	 */
+	raid10_find_phys(conf, r10_bio);
+	spin_lock_irq(&conf->device_lock);
+	for (i = 0;  i < conf->copies; i++) {
+		int d = r10_bio->devs[i].devnum;
+		if (conf->mirrors[d].rdev &&
+		    !conf->mirrors[d].rdev->faulty) {
+			atomic_inc(&conf->mirrors[d].rdev->nr_pending);
+			r10_bio->devs[i].bio = bio;
+		} else
+			r10_bio->devs[i].bio = NULL;
+	}
+	spin_unlock_irq(&conf->device_lock);
+
+	atomic_set(&r10_bio->remaining, 1);
+	md_write_start(mddev);
+	for (i = 0; i < conf->copies; i++) {
+		struct bio *mbio;
+		int d = r10_bio->devs[i].devnum;
+		if (!r10_bio->devs[i].bio)
+			continue;
+
+		mbio = bio_clone(bio, GFP_NOIO);
+		r10_bio->devs[i].bio = mbio;
+
+		mbio->bi_sector	= r10_bio->devs[i].addr+
+			conf->mirrors[d].rdev->data_offset;
+		mbio->bi_bdev = conf->mirrors[d].rdev->bdev;
+		mbio->bi_end_io	= raid10_end_write_request;
+		mbio->bi_rw = WRITE;
+		mbio->bi_private = r10_bio;
+
+		atomic_inc(&r10_bio->remaining);
+		generic_make_request(mbio);
+	}
+
+	if (atomic_dec_and_test(&r10_bio->remaining)) {
+		md_write_end(mddev);
+		raid_end_bio_io(r10_bio);
+	}
+
+	return 0;
+}
+
+static void status(struct seq_file *seq, mddev_t *mddev)
+{
+	conf_t *conf = mddev_to_conf(mddev);
+	int i;
+
+	if (conf->near_copies < conf->raid_disks)
+		seq_printf(seq, " %dK chunks", mddev->chunk_size/1024);
+	if (conf->near_copies > 1)
+		seq_printf(seq, " %d near-copies", conf->near_copies);
+	if (conf->far_copies > 1)
+		seq_printf(seq, " %d far-copies", conf->far_copies);
+
+	seq_printf(seq, " [%d/%d] [", conf->raid_disks,
+						conf->working_disks);
+	for (i = 0; i < conf->raid_disks; i++)
+		seq_printf(seq, "%s",
+			      conf->mirrors[i].rdev &&
+			      conf->mirrors[i].rdev->in_sync ? "U" : "_");
+	seq_printf(seq, "]");
+}
+
+static void error(mddev_t *mddev, mdk_rdev_t *rdev)
+{
+	char b[BDEVNAME_SIZE];
+	conf_t *conf = mddev_to_conf(mddev);
+
+	/*
+	 * If it is not operational, then we have already marked it as dead
+	 * else if it is the last working disks, ignore the error, let the
+	 * next level up know.
+	 * else mark the drive as failed
+	 */
+	if (rdev->in_sync
+	    && conf->working_disks == 1)
+		/*
+		 * Don't fail the drive, just return an IO error.
+		 * The test should really be more sophisticated than
+		 * "working_disks == 1", but it isn't critical, and
+		 * can wait until we do more sophisticated "is the drive
+		 * really dead" tests...
+		 */
+		return;
+	if (rdev->in_sync) {
+		mddev->degraded++;
+		conf->working_disks--;
+		/*
+		 * if recovery is running, make sure it aborts.
+		 */
+		set_bit(MD_RECOVERY_ERR, &mddev->recovery);
+	}
+	rdev->in_sync = 0;
+	rdev->faulty = 1;
+	mddev->sb_dirty = 1;
+	printk(KERN_ALERT "raid10: Disk failure on %s, disabling device. \n"
+		"	Operation continuing on %d devices\n",
+		bdevname(rdev->bdev,b), conf->working_disks);
+}
+
+static void print_conf(conf_t *conf)
+{
+	int i;
+	mirror_info_t *tmp;
+
+	printk("RAID10 conf printout:\n");
+	if (!conf) {
+		printk("(!conf)\n");
+		return;
+	}
+	printk(" --- wd:%d rd:%d\n", conf->working_disks,
+		conf->raid_disks);
+
+	for (i = 0; i < conf->raid_disks; i++) {
+		char b[BDEVNAME_SIZE];
+		tmp = conf->mirrors + i;
+		if (tmp->rdev)
+			printk(" disk %d, wo:%d, o:%d, dev:%s\n",
+				i, !tmp->rdev->in_sync, !tmp->rdev->faulty,
+				bdevname(tmp->rdev->bdev,b));
+	}
+}
+
+static void close_sync(conf_t *conf)
+{
+	spin_lock_irq(&conf->resync_lock);
+	wait_event_lock_irq(conf->wait_resume, !conf->barrier,
+			    conf->resync_lock, 	unplug_slaves(conf->mddev));
+	spin_unlock_irq(&conf->resync_lock);
+
+	if (conf->barrier) BUG();
+	if (waitqueue_active(&conf->wait_idle)) BUG();
+
+	mempool_destroy(conf->r10buf_pool);
+	conf->r10buf_pool = NULL;
+}
+
+static int raid10_spare_active(mddev_t *mddev)
+{
+	int i;
+	conf_t *conf = mddev->private;
+	mirror_info_t *tmp;
+
+	spin_lock_irq(&conf->device_lock);
+	/*
+	 * Find all non-in_sync disks within the RAID10 configuration
+	 * and mark them in_sync
+	 */
+	for (i = 0; i < conf->raid_disks; i++) {
+		tmp = conf->mirrors + i;
+		if (tmp->rdev
+		    && !tmp->rdev->faulty
+		    && !tmp->rdev->in_sync) {
+			conf->working_disks++;
+			mddev->degraded--;
+			tmp->rdev->in_sync = 1;
+		}
+	}
+	spin_unlock_irq(&conf->device_lock);
+
+	print_conf(conf);
+	return 0;
+}
+
+
+static int raid10_add_disk(mddev_t *mddev, mdk_rdev_t *rdev)
+{
+	conf_t *conf = mddev->private;
+	int found = 0;
+	int mirror;
+	mirror_info_t *p;
+
+	if (mddev->recovery_cp < MaxSector)
+		/* only hot-add to in-sync arrays, as recovery is
+		 * very different from resync
+		 */
+		return 0;
+	spin_lock_irq(&conf->device_lock);
+	for (mirror=0; mirror < mddev->raid_disks; mirror++)
+		if ( !(p=conf->mirrors+mirror)->rdev) {
+			p->rdev = rdev;
+
+			blk_queue_stack_limits(mddev->queue,
+					       rdev->bdev->bd_disk->queue);
+			/* as we don't honour merge_bvec_fn, we must never risk
+			 * violating it, so limit ->max_sector to one PAGE, as
+			 * a one page request is never in violation.
+			 */
+			if (rdev->bdev->bd_disk->queue->merge_bvec_fn &&
+			    mddev->queue->max_sectors > (PAGE_SIZE>>9))
+				mddev->queue->max_sectors = (PAGE_SIZE>>9);
+
+			p->head_position = 0;
+			rdev->raid_disk = mirror;
+			found = 1;
+			break;
+		}
+	spin_unlock_irq(&conf->device_lock);
+
+	print_conf(conf);
+	return found;
+}
+
+static int raid10_remove_disk(mddev_t *mddev, int number)
+{
+	conf_t *conf = mddev->private;
+	int err = 1;
+	mirror_info_t *p = conf->mirrors+ number;
+
+	print_conf(conf);
+	spin_lock_irq(&conf->device_lock);
+	if (p->rdev) {
+		if (p->rdev->in_sync ||
+		    atomic_read(&p->rdev->nr_pending)) {
+			err = -EBUSY;
+			goto abort;
+		}
+		p->rdev = NULL;
+		err = 0;
+	}
+	if (err)
+		MD_BUG();
+abort:
+	spin_unlock_irq(&conf->device_lock);
+
+	print_conf(conf);
+	return err;
+}
+
+
+static int end_sync_read(struct bio *bio, unsigned int bytes_done, int error)
+{
+	int uptodate = test_bit(BIO_UPTODATE, &bio->bi_flags);
+	r10bio_t * r10_bio = (r10bio_t *)(bio->bi_private);
+	conf_t *conf = mddev_to_conf(r10_bio->mddev);
+	int i,d;
+
+	if (bio->bi_size)
+		return 1;
+
+	for (i=0; i<conf->copies; i++)
+		if (r10_bio->devs[i].bio == bio)
+			break;
+	if (i == conf->copies)
+		BUG();
+	update_head_pos(i, r10_bio);
+	d = r10_bio->devs[i].devnum;
+	if (!uptodate)
+		md_error(r10_bio->mddev,
+			 conf->mirrors[d].rdev);
+
+	/* for reconstruct, we always reschedule after a read.
+	 * for resync, only after all reads 
+	 */
+	if (test_bit(R10BIO_IsRecover, &r10_bio->state) ||
+	    atomic_dec_and_test(&r10_bio->remaining)) {
+		/* we have read all the blocks,
+		 * do the comparison in process context in raid10d
+		 */
+		reschedule_retry(r10_bio);
+	}
+	rdev_dec_pending(conf->mirrors[d].rdev, conf->mddev);
+	return 0;
+}
+
+static int end_sync_write(struct bio *bio, unsigned int bytes_done, int error)
+{
+	int uptodate = test_bit(BIO_UPTODATE, &bio->bi_flags);
+	r10bio_t * r10_bio = (r10bio_t *)(bio->bi_private);
+	mddev_t *mddev = r10_bio->mddev;
+	conf_t *conf = mddev_to_conf(mddev);
+	int i,d;
+
+	if (bio->bi_size)
+		return 1;
+
+	for (i = 0; i < conf->copies; i++)
+		if (r10_bio->devs[i].bio == bio)
+			break;
+	d = r10_bio->devs[i].devnum;
+
+	if (!uptodate)
+		md_error(mddev, conf->mirrors[d].rdev);
+	update_head_pos(i, r10_bio);
+
+	while (atomic_dec_and_test(&r10_bio->remaining)) {
+		if (r10_bio->master_bio == NULL) {
+			/* the primary of several recovery bios */
+			md_done_sync(mddev, r10_bio->sectors, 1);
+			put_buf(r10_bio);
+			break;
+		} else {
+			r10bio_t *r10_bio2 = (r10bio_t *)r10_bio->master_bio;
+			put_buf(r10_bio);
+			r10_bio = r10_bio2;
+		}
+	}
+	rdev_dec_pending(conf->mirrors[d].rdev, mddev);
+	return 0;
+}
+
+/*
+ * Note: sync and recover and handled very differently for raid10
+ * This code is for resync.
+ * For resync, we read through virtual addresses and read all blocks.
+ * If there is any error, we schedule a write.  The lowest numbered
+ * drive is authoritative.
+ * However requests come for physical address, so we need to map.
+ * For every physical address there are raid_disks/copies virtual addresses,
+ * which is always are least one, but is not necessarly an integer.
+ * This means that a physical address can span multiple chunks, so we may
+ * have to submit multiple io requests for a single sync request.
+ */
+/*
+ * We check if all blocks are in-sync and only write to blocks that
+ * aren't in sync
+ */
+static void sync_request_write(mddev_t *mddev, r10bio_t *r10_bio)
+{
+	conf_t *conf = mddev_to_conf(mddev);
+	int i, first;
+	struct bio *tbio, *fbio;
+
+	atomic_set(&r10_bio->remaining, 1);
+
+	/* find the first device with a block */
+	for (i=0; i<conf->copies; i++)
+		if (test_bit(BIO_UPTODATE, &r10_bio->devs[i].bio->bi_flags))
+			break;
+
+	if (i == conf->copies)
+		goto done;
+
+	first = i;
+	fbio = r10_bio->devs[i].bio;
+
+	/* now find blocks with errors */
+	for (i=first+1 ; i < conf->copies ; i++) {
+		int vcnt, j, d;
+
+		if (!test_bit(BIO_UPTODATE, &r10_bio->devs[i].bio->bi_flags))
+			continue;
+		/* We know that the bi_io_vec layout is the same for
+		 * both 'first' and 'i', so we just compare them.
+		 * All vec entries are PAGE_SIZE;
+		 */
+		tbio = r10_bio->devs[i].bio;
+		vcnt = r10_bio->sectors >> (PAGE_SHIFT-9);
+		for (j = 0; j < vcnt; j++)
+			if (memcmp(page_address(fbio->bi_io_vec[j].bv_page),
+				   page_address(tbio->bi_io_vec[j].bv_page),
+				   PAGE_SIZE))
+				break;
+		if (j == vcnt)
+			continue;
+		/* Ok, we need to write this bio
+		 * First we need to fixup bv_offset, bv_len and
+		 * bi_vecs, as the read request might have corrupted these
+		 */
+		tbio->bi_vcnt = vcnt;
+		tbio->bi_size = r10_bio->sectors << 9;
+		tbio->bi_idx = 0;
+		tbio->bi_phys_segments = 0;
+		tbio->bi_hw_segments = 0;
+		tbio->bi_hw_front_size = 0;
+		tbio->bi_hw_back_size = 0;
+		tbio->bi_flags &= ~(BIO_POOL_MASK - 1);
+		tbio->bi_flags |= 1 << BIO_UPTODATE;
+		tbio->bi_next = NULL;
+		tbio->bi_rw = WRITE;
+		tbio->bi_private = r10_bio;
+		tbio->bi_sector = r10_bio->devs[i].addr;
+
+		for (j=0; j < vcnt ; j++) {
+			tbio->bi_io_vec[j].bv_offset = 0;
+			tbio->bi_io_vec[j].bv_len = PAGE_SIZE;
+		
+			memcpy(page_address(tbio->bi_io_vec[j].bv_page),
+			       page_address(fbio->bi_io_vec[j].bv_page),
+			       PAGE_SIZE);
+		}
+		tbio->bi_end_io = end_sync_write;
+
+		d = r10_bio->devs[i].devnum;
+		atomic_inc(&conf->mirrors[d].rdev->nr_pending);
+		atomic_inc(&r10_bio->remaining);
+		md_sync_acct(conf->mirrors[d].rdev->bdev, tbio->bi_size >> 9);
+
+		generic_make_request(tbio);
+	}
+
+done:
+	if (atomic_dec_and_test(&r10_bio->remaining)) {
+		md_done_sync(mddev, r10_bio->sectors, 1);
+		put_buf(r10_bio);
+	}
+}
+
+/*
+ * Now for the recovery code.
+ * Recovery happens across physical sectors.
+ * We recover all non-is_sync drives by finding the virtual address of
+ * each, and then choose a working drive that also has that virt address.
+ * There is a separate r10_bio for each non-in_sync drive.
+ * Only the first two slots are in use. The first for reading,
+ * The second for writing.
+ *
+ */
+
+static void recovery_request_write(mddev_t *mddev, r10bio_t *r10_bio)
+{
+	conf_t *conf = mddev_to_conf(mddev);
+	int i, d;
+	struct bio *bio, *wbio;
+
+
+	/* move the pages across to the second bio
+	 * and submit the write request
+	 */
+	bio = r10_bio->devs[0].bio;
+	wbio = r10_bio->devs[1].bio;
+	for (i=0; i < wbio->bi_vcnt; i++) {
+		struct page *p = bio->bi_io_vec[i].bv_page;
+		bio->bi_io_vec[i].bv_page = wbio->bi_io_vec[i].bv_page;
+		wbio->bi_io_vec[i].bv_page = p;
+	}
+	d = r10_bio->devs[1].devnum;
+
+	atomic_inc(&conf->mirrors[d].rdev->nr_pending);
+	md_sync_acct(conf->mirrors[d].rdev->bdev, wbio->bi_size >> 9);
+	generic_make_request(wbio);
+}
+
+
+/*
+ * This is a kernel thread which:
+ *
+ *	1.	Retries failed read operations on working mirrors.
+ *	2.	Updates the raid superblock when problems encounter.
+ *	3.	Performs writes following reads for array syncronising.
+ */
+
+static void raid10d(mddev_t *mddev)
+{
+	r10bio_t *r10_bio;
+	struct bio *bio;
+	unsigned long flags;
+	conf_t *conf = mddev_to_conf(mddev);
+	struct list_head *head = &conf->retry_list;
+	int unplug=0;
+	mdk_rdev_t *rdev;
+
+	md_check_recovery(mddev);
+	md_handle_safemode(mddev);
+
+	for (;;) {
+		char b[BDEVNAME_SIZE];
+		spin_lock_irqsave(&conf->device_lock, flags);
+		if (list_empty(head))
+			break;
+		r10_bio = list_entry(head->prev, r10bio_t, retry_list);
+		list_del(head->prev);
+		spin_unlock_irqrestore(&conf->device_lock, flags);
+
+		mddev = r10_bio->mddev;
+		conf = mddev_to_conf(mddev);
+		if (test_bit(R10BIO_IsSync, &r10_bio->state)) {
+			sync_request_write(mddev, r10_bio);
+			unplug = 1;
+		} else 	if (test_bit(R10BIO_IsRecover, &r10_bio->state)) {
+			recovery_request_write(mddev, r10_bio);
+			unplug = 1;
+		} else {
+			int mirror;
+			bio = r10_bio->devs[r10_bio->read_slot].bio;
+			r10_bio->devs[r10_bio->read_slot].bio = NULL;
+			mirror = read_balance(conf, r10_bio);
+			r10_bio->devs[r10_bio->read_slot].bio = bio;
+			if (mirror == -1) {
+				printk(KERN_ALERT "raid10: %s: unrecoverable I/O"
+				       " read error for block %llu\n",
+				       bdevname(bio->bi_bdev,b),
+				       (unsigned long long)r10_bio->sector);
+				raid_end_bio_io(r10_bio);
+			} else {
+				rdev = conf->mirrors[mirror].rdev;
+				if (printk_ratelimit())
+					printk(KERN_ERR "raid10: %s: redirecting sector %llu to"
+					       " another mirror\n",
+					       bdevname(rdev->bdev,b),
+					       (unsigned long long)r10_bio->sector);
+				bio->bi_bdev = rdev->bdev;
+				bio->bi_sector = r10_bio->devs[r10_bio->read_slot].addr
+					+ rdev->data_offset;
+				bio->bi_next = NULL;
+				bio->bi_flags &= (1<<BIO_CLONED);
+				bio->bi_flags |= 1 << BIO_UPTODATE;
+				bio->bi_idx = 0;
+				bio->bi_size = r10_bio->sectors << 9;
+				bio->bi_rw = READ;
+				unplug = 1;
+				generic_make_request(bio);
+			}
+		}
+	}
+	spin_unlock_irqrestore(&conf->device_lock, flags);
+	if (unplug)
+		unplug_slaves(mddev);
+}
+
+
+static int init_resync(conf_t *conf)
+{
+	int buffs;
+
+	buffs = RESYNC_WINDOW / RESYNC_BLOCK_SIZE;
+	if (conf->r10buf_pool)
+		BUG();
+	conf->r10buf_pool = mempool_create(buffs, r10buf_pool_alloc, r10buf_pool_free, conf);
+	if (!conf->r10buf_pool)
+		return -ENOMEM;
+	conf->next_resync = 0;
+	return 0;
+}
+
+/*
+ * perform a "sync" on one "block"
+ *
+ * We need to make sure that no normal I/O request - particularly write
+ * requests - conflict with active sync requests.
+ *
+ * This is achieved by tracking pending requests and a 'barrier' concept
+ * that can be installed to exclude normal IO requests.
+ *
+ * Resync and recovery are handled very differently.
+ * We differentiate by looking at MD_RECOVERY_SYNC in mddev->recovery.
+ *
+ * For resync, we iterate over virtual addresses, read all copies,
+ * and update if there are differences.  If only one copy is live,
+ * skip it.
+ * For recovery, we iterate over physical addresses, read a good
+ * value for each non-in_sync drive, and over-write.
+ *
+ * So, for recovery we may have several outstanding complex requests for a
+ * given address, one for each out-of-sync device.  We model this by allocating
+ * a number of r10_bio structures, one for each out-of-sync device.
+ * As we setup these structures, we collect all bio's together into a list
+ * which we then process collectively to add pages, and then process again
+ * to pass to generic_make_request.
+ *
+ * The r10_bio structures are linked using a borrowed master_bio pointer.
+ * This link is counted in ->remaining.  When the r10_bio that points to NULL 
+ * has its remaining count decremented to 0, the whole complex operation 
+ * is complete.
+ *
+ */
+
+static int sync_request(mddev_t *mddev, sector_t sector_nr, int go_faster)
+{
+	conf_t *conf = mddev_to_conf(mddev);
+	r10bio_t *r10_bio;
+	struct bio *biolist = NULL, *bio;
+	sector_t max_sector, nr_sectors;
+	int disk;
+	int i;
+
+	sector_t sectors_skipped = 0;
+	int chunks_skipped = 0;
+
+	if (!conf->r10buf_pool)
+		if (init_resync(conf))
+			return -ENOMEM;
+
+ skipped:
+	max_sector = mddev->size << 1;
+	if (test_bit(MD_RECOVERY_SYNC, &mddev->recovery))
+		max_sector = mddev->resync_max_sectors;
+	if (sector_nr >= max_sector) {
+		close_sync(conf);
+		return sectors_skipped;
+	}
+	if (chunks_skipped >= conf->raid_disks) {
+		/* if there has been nothing to do on any drive,
+		 * then there is nothing to do at all..
+		 */
+		sector_t sec = max_sector - sector_nr;
+		md_done_sync(mddev, sec, 1);
+		return sec + sectors_skipped;
+	}
+
+	/* make sure whole request will fit in a chunk - if chunks
+	 * are meaningful 
+	 */
+	if (conf->near_copies < conf->raid_disks &&
+	    max_sector > (sector_nr | conf->chunk_mask))
+		max_sector = (sector_nr | conf->chunk_mask) + 1;
+	/*
+	 * If there is non-resync activity waiting for us then
+	 * put in a delay to throttle resync.
+	 */
+	if (!go_faster && waitqueue_active(&conf->wait_resume))
+		schedule_timeout(HZ);
+	device_barrier(conf, sector_nr + RESYNC_SECTORS);
+
+	/* Again, very different code for resync and recovery.
+	 * Both must result in an r10bio with a list of bios that
+	 * have bi_end_io, bi_sector, bi_bdev set,
+	 * and bi_private set to the r10bio.
+	 * For recovery, we may actually create several r10bios
+	 * with 2 bios in each, that correspond to the bios in the main one.
+	 * In this case, the subordinate r10bios link back through a
+	 * borrowed master_bio pointer, and the counter in the master
+	 * includes a ref from each subordinate.
+	 */
+	/* First, we decide what to do and set ->bi_end_io
+	 * To end_sync_read if we want to read, and
+	 * end_sync_write if we will want to write.
+	 */
+
+	if (!test_bit(MD_RECOVERY_SYNC, &mddev->recovery)) {
+		/* recovery... the complicated one */
+		int i, j, k;
+		r10_bio = NULL;
+
+		for (i=0 ; i<conf->raid_disks; i++)
+			if (conf->mirrors[i].rdev &&
+			    !conf->mirrors[i].rdev->in_sync) {
+				/* want to reconstruct this device */
+				r10bio_t *rb2 = r10_bio;
+
+				r10_bio = mempool_alloc(conf->r10buf_pool, GFP_NOIO);
+				spin_lock_irq(&conf->resync_lock);
+				conf->nr_pending++;
+				if (rb2) conf->barrier++;
+				spin_unlock_irq(&conf->resync_lock);
+				atomic_set(&r10_bio->remaining, 0);
+
+				r10_bio->master_bio = (struct bio*)rb2;
+				if (rb2)
+					atomic_inc(&rb2->remaining);
+				r10_bio->mddev = mddev;
+				set_bit(R10BIO_IsRecover, &r10_bio->state);
+				r10_bio->sector = raid10_find_virt(conf, sector_nr, i);
+				raid10_find_phys(conf, r10_bio);
+				for (j=0; j<conf->copies;j++) {
+					int d = r10_bio->devs[j].devnum;
+					if (conf->mirrors[d].rdev &&
+					    conf->mirrors[d].rdev->in_sync) {
+						/* This is where we read from */
+						bio = r10_bio->devs[0].bio;
+						bio->bi_next = biolist;
+						biolist = bio;
+						bio->bi_private = r10_bio;
+						bio->bi_end_io = end_sync_read;
+						bio->bi_rw = 0;
+						bio->bi_sector = r10_bio->devs[j].addr +
+							conf->mirrors[d].rdev->data_offset;
+						bio->bi_bdev = conf->mirrors[d].rdev->bdev;
+						atomic_inc(&conf->mirrors[d].rdev->nr_pending);
+						atomic_inc(&r10_bio->remaining);
+						/* and we write to 'i' */
+
+						for (k=0; k<conf->copies; k++)
+							if (r10_bio->devs[k].devnum == i)
+								break;
+						bio = r10_bio->devs[1].bio;
+						bio->bi_next = biolist;
+						biolist = bio;
+						bio->bi_private = r10_bio;
+						bio->bi_end_io = end_sync_write;
+						bio->bi_rw = 1;
+						bio->bi_sector = r10_bio->devs[k].addr +
+							conf->mirrors[i].rdev->data_offset;
+						bio->bi_bdev = conf->mirrors[i].rdev->bdev;
+
+						r10_bio->devs[0].devnum = d;
+						r10_bio->devs[1].devnum = i;
+
+						break;
+					}
+				}
+				if (j == conf->copies) {
+					BUG();
+				}
+			}
+		if (biolist == NULL) {
+			while (r10_bio) {
+				r10bio_t *rb2 = r10_bio;
+				r10_bio = (r10bio_t*) rb2->master_bio;
+				rb2->master_bio = NULL;
+				put_buf(rb2);
+			}
+			goto giveup;
+		}
+	} else {
+		/* resync. Schedule a read for every block at this virt offset */
+		int count = 0;
+		r10_bio = mempool_alloc(conf->r10buf_pool, GFP_NOIO);
+
+		spin_lock_irq(&conf->resync_lock);
+		conf->nr_pending++;
+		spin_unlock_irq(&conf->resync_lock);
+
+		r10_bio->mddev = mddev;
+		atomic_set(&r10_bio->remaining, 0);
+
+		r10_bio->master_bio = NULL;
+		r10_bio->sector = sector_nr;
+		set_bit(R10BIO_IsSync, &r10_bio->state);
+		raid10_find_phys(conf, r10_bio);
+		r10_bio->sectors = (sector_nr | conf->chunk_mask) - sector_nr +1;
+		spin_lock_irq(&conf->device_lock);
+		for (i=0; i<conf->copies; i++) {
+			int d = r10_bio->devs[i].devnum;
+			bio = r10_bio->devs[i].bio;
+			bio->bi_end_io = NULL;
+			if (conf->mirrors[d].rdev == NULL ||
+			    conf->mirrors[d].rdev->faulty)
+				continue;
+			atomic_inc(&conf->mirrors[d].rdev->nr_pending);
+			atomic_inc(&r10_bio->remaining);
+			bio->bi_next = biolist;
+			biolist = bio;
+			bio->bi_private = r10_bio;
+			bio->bi_end_io = end_sync_read;
+			bio->bi_rw = 0;
+			bio->bi_sector = r10_bio->devs[i].addr +
+				conf->mirrors[d].rdev->data_offset;
+			bio->bi_bdev = conf->mirrors[d].rdev->bdev;
+			count++;
+		}
+		spin_unlock_irq(&conf->device_lock);
+		if (count < 2) {
+			for (i=0; i<conf->copies; i++) {
+				int d = r10_bio->devs[i].devnum;
+				if (r10_bio->devs[i].bio->bi_end_io)
+					atomic_dec(&conf->mirrors[d].rdev->nr_pending);
+			}
+			put_buf(r10_bio);
+			goto giveup;
+		}
+	}
+
+	for (bio = biolist; bio ; bio=bio->bi_next) {
+
+		bio->bi_flags &= ~(BIO_POOL_MASK - 1);
+		if (bio->bi_end_io)
+			bio->bi_flags |= 1 << BIO_UPTODATE;
+		bio->bi_vcnt = 0;
+		bio->bi_idx = 0;
+		bio->bi_phys_segments = 0;
+		bio->bi_hw_segments = 0;
+		bio->bi_size = 0;
+	}
+
+	nr_sectors = 0;
+	do {
+		struct page *page;
+		int len = PAGE_SIZE;
+		disk = 0;
+		if (sector_nr + (len>>9) > max_sector)
+			len = (max_sector - sector_nr) << 9;
+		if (len == 0)
+			break;
+		for (bio= biolist ; bio ; bio=bio->bi_next) {
+			page = bio->bi_io_vec[bio->bi_vcnt].bv_page;
+			if (bio_add_page(bio, page, len, 0) == 0) {
+				/* stop here */
+				struct bio *bio2;
+				bio->bi_io_vec[bio->bi_vcnt].bv_page = page;
+				for (bio2 = biolist; bio2 && bio2 != bio; bio2 = bio2->bi_next) {
+					/* remove last page from this bio */
+					bio2->bi_vcnt--;
+					bio2->bi_size -= len;
+					bio2->bi_flags &= ~(1<< BIO_SEG_VALID);
+				}
+				goto bio_full;
+			}
+			disk = i;
+		}
+		nr_sectors += len>>9;
+		sector_nr += len>>9;
+	} while (biolist->bi_vcnt < RESYNC_PAGES);
+ bio_full:
+	r10_bio->sectors = nr_sectors;
+
+	while (biolist) {
+		bio = biolist;
+		biolist = biolist->bi_next;
+
+		bio->bi_next = NULL;
+		r10_bio = bio->bi_private;
+		r10_bio->sectors = nr_sectors;
+
+		if (bio->bi_end_io == end_sync_read) {
+			md_sync_acct(bio->bi_bdev, nr_sectors);
+			generic_make_request(bio);
+		}
+	}
+
+	return nr_sectors;
+ giveup:
+	/* There is nowhere to write, so all non-sync
+	 * drives must be failed, so try the next chunk...
+	 */
+	{
+	int sec = max_sector - sector_nr;
+	sectors_skipped += sec;
+	chunks_skipped ++;
+	sector_nr = max_sector;
+	md_done_sync(mddev, sec, 1);
+	goto skipped;
+	}
+}
+
+static int run(mddev_t *mddev)
+{
+	conf_t *conf;
+	int i, disk_idx;
+	mirror_info_t *disk;
+	mdk_rdev_t *rdev;
+	struct list_head *tmp;
+	int nc, fc;
+	sector_t stride, size;
+
+	if (mddev->level != 10) {
+		printk(KERN_ERR "raid10: %s: raid level not set correctly... (%d)\n",
+		       mdname(mddev), mddev->level);
+		goto out;
+	}
+	nc = mddev->layout & 255;
+	fc = (mddev->layout >> 8) & 255;
+	if ((nc*fc) <2 || (nc*fc) > mddev->raid_disks ||
+	    (mddev->layout >> 16)) {
+		printk(KERN_ERR "raid10: %s: unsupported raid10 layout: 0x%8x\n",
+		       mdname(mddev), mddev->layout);
+		goto out;
+	}
+	/*
+	 * copy the already verified devices into our private RAID10
+	 * bookkeeping area. [whatever we allocate in run(),
+	 * should be freed in stop()]
+	 */
+	conf = kmalloc(sizeof(conf_t), GFP_KERNEL);
+	mddev->private = conf;
+	if (!conf) {
+		printk(KERN_ERR "raid10: couldn't allocate memory for %s\n",
+			mdname(mddev));
+		goto out;
+	}
+	memset(conf, 0, sizeof(*conf));
+	conf->mirrors = kmalloc(sizeof(struct mirror_info)*mddev->raid_disks,
+				 GFP_KERNEL);
+	if (!conf->mirrors) {
+		printk(KERN_ERR "raid10: couldn't allocate memory for %s\n",
+		       mdname(mddev));
+		goto out_free_conf;
+	}
+	memset(conf->mirrors, 0, sizeof(struct mirror_info)*mddev->raid_disks);
+
+	conf->near_copies = nc;
+	conf->far_copies = fc;
+	conf->copies = nc*fc;
+	conf->chunk_mask = (sector_t)(mddev->chunk_size>>9)-1;
+	conf->chunk_shift = ffz(~mddev->chunk_size) - 9;
+	stride = mddev->size >> (conf->chunk_shift-1);
+	sector_div(stride, fc);
+	conf->stride = stride << conf->chunk_shift;
+
+	conf->r10bio_pool = mempool_create(NR_RAID10_BIOS, r10bio_pool_alloc,
+						r10bio_pool_free, conf);
+	if (!conf->r10bio_pool) {
+		printk(KERN_ERR "raid10: couldn't allocate memory for %s\n",
+			mdname(mddev));
+		goto out_free_conf;
+	}
+	mddev->queue->unplug_fn = raid10_unplug;
+
+	mddev->queue->issue_flush_fn = raid10_issue_flush;
+
+	ITERATE_RDEV(mddev, rdev, tmp) {
+		disk_idx = rdev->raid_disk;
+		if (disk_idx >= mddev->raid_disks
+		    || disk_idx < 0)
+			continue;
+		disk = conf->mirrors + disk_idx;
+
+		disk->rdev = rdev;
+
+		blk_queue_stack_limits(mddev->queue,
+				       rdev->bdev->bd_disk->queue);
+		/* as we don't honour merge_bvec_fn, we must never risk
+		 * violating it, so limit ->max_sector to one PAGE, as
+		 * a one page request is never in violation.
+		 */
+		if (rdev->bdev->bd_disk->queue->merge_bvec_fn &&
+		    mddev->queue->max_sectors > (PAGE_SIZE>>9))
+			mddev->queue->max_sectors = (PAGE_SIZE>>9);
+
+		disk->head_position = 0;
+		if (!rdev->faulty && rdev->in_sync)
+			conf->working_disks++;
+	}
+	conf->raid_disks = mddev->raid_disks;
+	conf->mddev = mddev;
+	conf->device_lock = SPIN_LOCK_UNLOCKED;
+	INIT_LIST_HEAD(&conf->retry_list);
+
+	conf->resync_lock = SPIN_LOCK_UNLOCKED;
+	init_waitqueue_head(&conf->wait_idle);
+	init_waitqueue_head(&conf->wait_resume);
+
+	if (!conf->working_disks) {
+		printk(KERN_ERR "raid10: no operational mirrors for %s\n",
+			mdname(mddev));
+		goto out_free_conf;
+	}
+
+	mddev->degraded = 0;
+	for (i = 0; i < conf->raid_disks; i++) {
+
+		disk = conf->mirrors + i;
+
+		if (!disk->rdev) {
+			disk->head_position = 0;
+			mddev->degraded++;
+		}
+	}
+
+
+	mddev->thread = md_register_thread(raid10d, mddev, "%s_raid10");
+	if (!mddev->thread) {
+		printk(KERN_ERR
+		       "raid10: couldn't allocate thread for %s\n",
+		       mdname(mddev));
+		goto out_free_conf;
+	}
+
+	printk(KERN_INFO
+		"raid10: raid set %s active with %d out of %d devices\n",
+		mdname(mddev), mddev->raid_disks - mddev->degraded,
+		mddev->raid_disks);
+	/*
+	 * Ok, everything is just fine now
+	 */
+	size = conf->stride * conf->raid_disks;
+	sector_div(size, conf->near_copies);
+	mddev->array_size = size/2;
+	mddev->resync_max_sectors = size;
+
+	/* Calculate max read-ahead size.
+	 * We need to readahead at least twice a whole stripe....
+	 * maybe...
+	 */
+	{
+		int stripe = conf->raid_disks * mddev->chunk_size / PAGE_CACHE_SIZE;
+		stripe /= conf->near_copies;
+		if (mddev->queue->backing_dev_info.ra_pages < 2* stripe)
+			mddev->queue->backing_dev_info.ra_pages = 2* stripe;
+	}
+
+	if (conf->near_copies < mddev->raid_disks)
+		blk_queue_merge_bvec(mddev->queue, raid10_mergeable_bvec);
+	return 0;
+
+out_free_conf:
+	if (conf->r10bio_pool)
+		mempool_destroy(conf->r10bio_pool);
+	if (conf->mirrors)
+		kfree(conf->mirrors);
+	kfree(conf);
+	mddev->private = NULL;
+out:
+	return -EIO;
+}
+
+static int stop(mddev_t *mddev)
+{
+	conf_t *conf = mddev_to_conf(mddev);
+
+	md_unregister_thread(mddev->thread);
+	mddev->thread = NULL;
+	if (conf->r10bio_pool)
+		mempool_destroy(conf->r10bio_pool);
+	if (conf->mirrors)
+		kfree(conf->mirrors);
+	kfree(conf);
+	mddev->private = NULL;
+	return 0;
+}
+
+
+static mdk_personality_t raid10_personality =
+{
+	.name		= "raid10",
+	.owner		= THIS_MODULE,
+	.make_request	= make_request,
+	.run		= run,
+	.stop		= stop,
+	.status		= status,
+	.error_handler	= error,
+	.hot_add_disk	= raid10_add_disk,
+	.hot_remove_disk= raid10_remove_disk,
+	.spare_active	= raid10_spare_active,
+	.sync_request	= sync_request,
+};
+
+static int __init raid_init(void)
+{
+	return register_md_personality(RAID10, &raid10_personality);
+}
+
+static void raid_exit(void)
+{
+	unregister_md_personality(RAID10);
+}
+
+module_init(raid_init);
+module_exit(raid_exit);
+MODULE_LICENSE("GPL");
+MODULE_ALIAS("md-personality-9"); /* RAID10 */

diff ./include/linux/raid/md_k.h~current~ ./include/linux/raid/md_k.h
--- ./include/linux/raid/md_k.h~current~	2004-08-23 13:06:02.000000000 +1000
+++ ./include/linux/raid/md_k.h	2004-08-23 12:58:29.000000000 +1000
@@ -24,7 +24,8 @@
 #define HSM               6UL
 #define MULTIPATH         7UL
 #define RAID6		  8UL
-#define MAX_PERSONALITY   9UL
+#define	RAID10		  9UL
+#define MAX_PERSONALITY   10UL
 
 #define	LEVEL_MULTIPATH		(-4)
 #define	LEVEL_LINEAR		(-1)
@@ -43,6 +44,7 @@ static inline int pers_to_level (int per
 		case RAID1:		return 1;
 		case RAID5:		return 5;
 		case RAID6:		return 6;
+		case RAID10:		return 10;
 	}
 	BUG();
 	return MD_RESERVED;
@@ -60,6 +62,7 @@ static inline int level_to_pers (int lev
 		case 4:
 		case 5: return RAID5;
 		case 6: return RAID6;
+		case 10: return RAID10;
 	}
 	return MD_RESERVED;
 }

diff ./include/linux/raid/raid10.h~current~ ./include/linux/raid/raid10.h
--- ./include/linux/raid/raid10.h~current~	2004-08-23 13:06:02.000000000 +1000
+++ ./include/linux/raid/raid10.h	2004-08-23 12:49:27.000000000 +1000
@@ -0,0 +1,103 @@
+#ifndef _RAID10_H
+#define _RAID10_H
+
+#include <linux/raid/md.h>
+
+typedef struct mirror_info mirror_info_t;
+
+struct mirror_info {
+	mdk_rdev_t	*rdev;
+	sector_t	head_position;
+};
+
+typedef struct r10bio_s r10bio_t;
+
+struct r10_private_data_s {
+	mddev_t			*mddev;
+	mirror_info_t		*mirrors;
+	int			raid_disks;
+	int			working_disks;
+	spinlock_t		device_lock;
+
+	/* geometry */
+	int			near_copies;  /* number of copies layed out raid0 style */
+	int 			far_copies;   /* number of copies layed out
+					       * at large strides across drives
+					       */
+	int			copies;	      /* near_copies * far_copies.
+					       * must be <= raid_disks
+					       */
+	sector_t		stride;	      /* distance between far copies.
+					       * This is size / far_copies
+					       */
+
+	int chunk_shift; /* shift from chunks to sectors */
+	sector_t chunk_mask;
+
+	struct list_head	retry_list;
+	/* for use when syncing mirrors: */
+
+	spinlock_t		resync_lock;
+	int nr_pending;
+	int barrier;
+	sector_t		next_resync;
+
+	wait_queue_head_t	wait_idle;
+	wait_queue_head_t	wait_resume;
+
+	mempool_t *r10bio_pool;
+	mempool_t *r10buf_pool;
+};
+
+typedef struct r10_private_data_s conf_t;
+
+/*
+ * this is the only point in the RAID code where we violate
+ * C type safety. mddev->private is an 'opaque' pointer.
+ */
+#define mddev_to_conf(mddev) ((conf_t *) mddev->private)
+
+/*
+ * this is our 'private' RAID10 bio.
+ *
+ * it contains information about what kind of IO operations were started
+ * for this RAID10 operation, and about their status:
+ */
+
+struct r10bio_s {
+	atomic_t		remaining; /* 'have we finished' count,
+					    * used from IRQ handlers
+					    */
+	sector_t		sector;	/* virtual sector number */
+	int			sectors;
+	unsigned long		state;
+	mddev_t			*mddev;
+	/*
+	 * original bio going to /dev/mdx
+	 */
+	struct bio		*master_bio;
+	/*
+	 * if the IO is in READ direction, then this is where we read
+	 */
+	int			read_slot;
+
+	struct list_head	retry_list;
+	/*
+	 * if the IO is in WRITE direction, then multiple bios are used,
+	 * one for each copy.
+	 * When resyncing we also use one for each copy.
+	 * When reconstructing, we use 2 bios, one for read, one for write.
+	 * We choose the number when they are allocated.
+	 */
+	struct {
+		struct bio		*bio;
+		sector_t addr;
+		int devnum;
+	} devs[0];
+};
+
+/* bits for r10bio.state */
+#define	R10BIO_Uptodate	0
+#define	R10BIO_IsSync	1
+#define	R10BIO_IsRecover 2
+#endif

^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 1 of 4] Assorted fixes/improvemnet to generic md resync code.
  2004-08-23  3:10 [PATCH md 0 of 4] Introduction NeilBrown
@ 2004-08-23  3:10 ` NeilBrown
  2004-08-23  3:10 ` [PATCH md 4 of 4] RAID10 module for MD NeilBrown
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 16+ messages in thread
From: NeilBrown @ 2004-08-23  3:10 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid


1/ Introduce "mddev->resync_max_sectors" so that an md personality
can ask for resync to cover a different address range than that of a
single drive.  raid10 will use this.

2/ fix is_mddev_idle so that if there seem to be a negative number
 of events, it doesn't immediately assume activity.

3/ make "sync_io" (the count of IO sectors used for array resync)
 an atomic_t to avoid SMP races.  

4/ Pass md_sync_acct a "block_device" rather than the containing "rdev",
  as the whole rdev isn't needed. Also make this an inline function.

5/ Make sure recovery gets interrupted on any error.

diff ./drivers/md/md.c~current~ ./drivers/md/md.c
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au>

### Diffstat output
 ./drivers/md/md.c           |   32 ++++++++++++++++++++++----------
 ./drivers/md/raid1.c        |    4 ++--
 ./drivers/md/raid5.c        |    5 +++--
 ./drivers/md/raid6main.c    |    5 +++--
 ./include/linux/genhd.h     |    2 +-
 ./include/linux/raid/md.h   |    1 -
 ./include/linux/raid/md_k.h |    6 ++++++
 7 files changed, 37 insertions(+), 18 deletions(-)

diff ./drivers/md/md.c~current~ ./drivers/md/md.c
--- ./drivers/md/md.c~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./drivers/md/md.c	2004-08-23 12:23:14.000000000 +1000
@@ -1684,6 +1684,8 @@ static int do_md_run(mddev_t * mddev)
 	mddev->pers = pers[pnum];
 	spin_unlock(&pers_lock);
 
+	mddev->resync_max_sectors = mddev->size << 1; /* may be over-ridden by personality */
+
 	err = mddev->pers->run(mddev);
 	if (err) {
 		printk(KERN_ERR "md: pers->run() failed ...\n");
@@ -2989,6 +2991,7 @@ void md_error(mddev_t *mddev, mdk_rdev_t
 	if (!mddev->pers->error_handler)
 		return;
 	mddev->pers->error_handler(mddev,rdev);
+	set_bit(MD_RECOVERY_INTR, &mddev->recovery);
 	set_bit(MD_RECOVERY_NEEDED, &mddev->recovery);
 	md_wakeup_thread(mddev->thread);
 }
@@ -3021,7 +3024,11 @@ static void status_resync(struct seq_fil
 	unsigned long max_blocks, resync, res, dt, db, rt;
 
 	resync = (mddev->curr_resync - atomic_read(&mddev->recovery_active))/2;
-	max_blocks = mddev->size;
+
+	if (test_bit(MD_RECOVERY_SYNC, &mddev->recovery))
+		max_blocks = mddev->resync_max_sectors >> 1;
+	else
+		max_blocks = mddev->size;
 
 	/*
 	 * Should not happen.
@@ -3257,11 +3264,6 @@ int unregister_md_personality(int pnum)
 	return 0;
 }
 
-void md_sync_acct(mdk_rdev_t *rdev, unsigned long nr_sectors)
-{
-	rdev->bdev->bd_contains->bd_disk->sync_io += nr_sectors;
-}
-
 static int is_mddev_idle(mddev_t *mddev)
 {
 	mdk_rdev_t * rdev;
@@ -3274,8 +3276,12 @@ static int is_mddev_idle(mddev_t *mddev)
 		struct gendisk *disk = rdev->bdev->bd_contains->bd_disk;
 		curr_events = disk_stat_read(disk, read_sectors) + 
 				disk_stat_read(disk, write_sectors) - 
-				disk->sync_io;
-		if ((curr_events - rdev->last_events) > 32) {
+				atomic_read(&disk->sync_io);
+		/* Allow some slack between valud of curr_events and last_events,
+		 * as there are some uninteresting races.
+		 * Note: the following is an unsigned comparison.
+		 */
+		if ((curr_events - rdev->last_events + 32) > 64) {
 			rdev->last_events = curr_events;
 			idle = 0;
 		}
@@ -3409,7 +3415,14 @@ static void md_do_sync(mddev_t *mddev)
 		}
 	} while (mddev->curr_resync < 2);
 
-	max_sectors = mddev->size << 1;
+	if (test_bit(MD_RECOVERY_SYNC, &mddev->recovery))
+		/* resync follows the size requested by the personality,
+		 * which default to physical size, but can be virtual size
+		 */
+		max_sectors = mddev->resync_max_sectors;
+	else
+		/* recovery follows the physical size of devices */
+		max_sectors = mddev->size << 1;
 
 	printk(KERN_INFO "md: syncing RAID array %s\n", mdname(mddev));
 	printk(KERN_INFO "md: minimum _guaranteed_ reconstruction speed:"
@@ -3832,7 +3845,6 @@ module_exit(md_exit)
 EXPORT_SYMBOL(register_md_personality);
 EXPORT_SYMBOL(unregister_md_personality);
 EXPORT_SYMBOL(md_error);
-EXPORT_SYMBOL(md_sync_acct);
 EXPORT_SYMBOL(md_done_sync);
 EXPORT_SYMBOL(md_write_start);
 EXPORT_SYMBOL(md_write_end);

diff ./drivers/md/raid1.c~current~ ./drivers/md/raid1.c
--- ./drivers/md/raid1.c~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./drivers/md/raid1.c	2004-08-23 12:20:59.000000000 +1000
@@ -903,7 +903,7 @@ static void sync_request_write(mddev_t *
 
 		atomic_inc(&conf->mirrors[i].rdev->nr_pending);
 		atomic_inc(&r1_bio->remaining);
-		md_sync_acct(conf->mirrors[i].rdev, wbio->bi_size >> 9);
+		md_sync_acct(conf->mirrors[i].rdev->bdev, wbio->bi_size >> 9);
 		generic_make_request(wbio);
 	}
 
@@ -1143,7 +1143,7 @@ static int sync_request(mddev_t *mddev, 
 	bio = r1_bio->bios[disk];
 	r1_bio->sectors = nr_sectors;
 
-	md_sync_acct(mirror->rdev, nr_sectors);
+	md_sync_acct(mirror->rdev->bdev, nr_sectors);
 
 	generic_make_request(bio);
 

diff ./drivers/md/raid5.c~current~ ./drivers/md/raid5.c
--- ./drivers/md/raid5.c~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./drivers/md/raid5.c	2004-08-23 12:20:59.000000000 +1000
@@ -1071,7 +1071,8 @@ static void handle_stripe(struct stripe_
 					PRINTK("Reading block %d (sync=%d)\n", 
 						i, syncing);
 					if (syncing)
-						md_sync_acct(conf->disks[i].rdev, STRIPE_SECTORS);
+						md_sync_acct(conf->disks[i].rdev->bdev,
+							     STRIPE_SECTORS);
 				}
 			}
 		}
@@ -1256,7 +1257,7 @@ static void handle_stripe(struct stripe_
  
 		if (rdev) {
 			if (test_bit(R5_Syncio, &sh->dev[i].flags))
-				md_sync_acct(rdev, STRIPE_SECTORS);
+				md_sync_acct(rdev->bdev, STRIPE_SECTORS);
 
 			bi->bi_bdev = rdev->bdev;
 			PRINTK("for %llu schedule op %ld on disc %d\n",

diff ./drivers/md/raid6main.c~current~ ./drivers/md/raid6main.c
--- ./drivers/md/raid6main.c~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./drivers/md/raid6main.c	2004-08-23 12:20:59.000000000 +1000
@@ -1208,7 +1208,8 @@ static void handle_stripe(struct stripe_
 					PRINTK("Reading block %d (sync=%d)\n",
 						i, syncing);
 					if (syncing)
-						md_sync_acct(conf->disks[i].rdev, STRIPE_SECTORS);
+						md_sync_acct(conf->disks[i].rdev->bdev,
+							     STRIPE_SECTORS);
 				}
 			}
 		}
@@ -1418,7 +1419,7 @@ static void handle_stripe(struct stripe_
 
 		if (rdev) {
 			if (test_bit(R5_Syncio, &sh->dev[i].flags))
-				md_sync_acct(rdev, STRIPE_SECTORS);
+				md_sync_acct(rdev->bdev, STRIPE_SECTORS);
 
 			bi->bi_bdev = rdev->bdev;
 			PRINTK("for %llu schedule op %ld on disc %d\n",

diff ./include/linux/genhd.h~current~ ./include/linux/genhd.h
--- ./include/linux/genhd.h~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./include/linux/genhd.h	2004-08-23 12:20:59.000000000 +1000
@@ -100,7 +100,7 @@ struct gendisk {
 	struct timer_rand_state *random;
 	int policy;
 
-	unsigned sync_io;		/* RAID */
+	atomic_t sync_io;		/* RAID */
 	unsigned long stamp, stamp_idle;
 	int in_flight;
 #ifdef	CONFIG_SMP

diff ./include/linux/raid/md.h~current~ ./include/linux/raid/md.h
--- ./include/linux/raid/md.h~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./include/linux/raid/md.h	2004-08-23 12:20:59.000000000 +1000
@@ -74,7 +74,6 @@ extern void md_write_start(mddev_t *mdde
 extern void md_write_end(mddev_t *mddev);
 extern void md_handle_safemode(mddev_t *mddev);
 extern void md_done_sync(mddev_t *mddev, int blocks, int ok);
-extern void md_sync_acct(mdk_rdev_t *rdev, unsigned long nr_sectors);
 extern void md_error (mddev_t *mddev, mdk_rdev_t *rdev);
 extern void md_unplug_mddev(mddev_t *mddev);
 

diff ./include/linux/raid/md_k.h~current~ ./include/linux/raid/md_k.h
--- ./include/linux/raid/md_k.h~current~	2004-08-23 12:07:44.000000000 +1000
+++ ./include/linux/raid/md_k.h	2004-08-23 12:20:59.000000000 +1000
@@ -216,6 +216,7 @@ struct mddev_s
 	unsigned long			resync_mark;	/* a recent timestamp */
 	sector_t			resync_mark_cnt;/* blocks written at resync_mark */
 
+	sector_t			resync_max_sectors; /* may be set by personality */
 	/* recovery/resync flags 
 	 * NEEDED:   we might need to start a resync/recover
 	 * RUNNING:  a thread is running, or about to be started
@@ -263,6 +264,11 @@ static inline void rdev_dec_pending(mdk_
 		set_bit(MD_RECOVERY_NEEDED, &mddev->recovery);
 }
 
+static inline void md_sync_acct(struct block_device *bdev, unsigned long nr_sectors)
+{
+        atomic_add(nr_sectors, &bdev->bd_contains->bd_disk->sync_io);
+}
+
 struct mdk_personality_s
 {
 	char *name;

^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 3 of 4] Remove most calls to __bdevname from md.c
  2004-08-23  3:10 [PATCH md 0 of 4] Introduction NeilBrown
  2004-08-23  3:10 ` [PATCH md 1 of 4] Assorted fixes/improvemnet to generic md resync code NeilBrown
  2004-08-23  3:10 ` [PATCH md 4 of 4] RAID10 module for MD NeilBrown
@ 2004-08-23  3:10 ` NeilBrown
  2004-08-23  3:10 ` [PATCH md 2 of 4] Assorted minor md/raid1 fixes NeilBrown
  3 siblings, 0 replies; 16+ messages in thread
From: NeilBrown @ 2004-08-23  3:10 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid


__bdevname now only prints major/minor number which isn't
much help.  So remove most calls to it from md.c, replacing
those that are useful by calls to bdevname (often printing the
message when the error is first detected rather than higher
up the call tree).

Also discard hot_generate_error which doesn't do anything
useful and never has.

Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au>

### Diffstat output
 ./drivers/md/md.c |   96 +++++++++++-------------------------------------------
 1 files changed, 20 insertions(+), 76 deletions(-)

diff ./drivers/md/md.c~current~ ./drivers/md/md.c
--- ./drivers/md/md.c~current~	2004-08-23 12:28:43.000000000 +1000
+++ ./drivers/md/md.c	2004-08-23 12:41:57.000000000 +1000
@@ -1111,20 +1111,24 @@ static void unbind_rdev_from_array(mdk_r
 /*
  * prevent the device from being mounted, repartitioned or
  * otherwise reused by a RAID array (or any other kernel
- * subsystem), by opening the device. [simply getting an
- * inode is not enough, the SCSI module usage code needs
- * an explicit open() on the device]
+ * subsystem), by bd_claiming the device.
  */
 static int lock_rdev(mdk_rdev_t *rdev, dev_t dev)
 {
 	int err = 0;
 	struct block_device *bdev;
+	char b[BDEVNAME_SIZE];
 
 	bdev = open_by_devnum(dev, FMODE_READ|FMODE_WRITE);
-	if (IS_ERR(bdev))
+	if (IS_ERR(bdev)) {
+		printk(KERN_ERR "md: could not open %s.\n",
+			__bdevname(dev, b));
 		return PTR_ERR(bdev);
+	}
 	err = bd_claim(bdev, rdev);
 	if (err) {
+		printk(KERN_ERR "md: could not bd_claim %s.\n",
+			bdevname(bdev, b));
 		blkdev_put(bdev);
 		return err;
 	}
@@ -1186,10 +1190,7 @@ static void export_array(mddev_t *mddev)
 
 static void print_desc(mdp_disk_t *desc)
 {
-	char b[BDEVNAME_SIZE];
-
-	printk(" DISK<N:%d,%s(%d,%d),R:%d,S:%d>\n", desc->number,
-		__bdevname(MKDEV(desc->major, desc->minor), b),
+	printk(" DISK<N:%d,(%d,%d),R:%d,S:%d>\n", desc->number,
 		desc->major,desc->minor,desc->raid_disk,desc->state);
 }
 
@@ -1381,8 +1382,7 @@ static mdk_rdev_t *md_import_device(dev_
 
 	rdev = (mdk_rdev_t *) kmalloc(sizeof(*rdev), GFP_KERNEL);
 	if (!rdev) {
-		printk(KERN_ERR "md: could not alloc mem for %s!\n", 
-			__bdevname(newdev, b));
+		printk(KERN_ERR "md: could not alloc mem for new device!\n");
 		return ERR_PTR(-ENOMEM);
 	}
 	memset(rdev, 0, sizeof(*rdev));
@@ -1391,11 +1391,9 @@ static mdk_rdev_t *md_import_device(dev_
 		goto abort_free;
 
 	err = lock_rdev(rdev, newdev);
-	if (err) {
-		printk(KERN_ERR "md: could not lock %s.\n",
-			__bdevname(newdev, b));
+	if (err)
 		goto abort_free;
-	}
+
 	rdev->desc_nr = -1;
 	rdev->faulty = 0;
 	rdev->in_sync = 0;
@@ -1953,11 +1951,9 @@ static int autostart_array(dev_t startde
 	mdk_rdev_t *start_rdev = NULL, *rdev;
 
 	start_rdev = md_import_device(startdev, 0, 0);
-	if (IS_ERR(start_rdev)) {
-		printk(KERN_WARNING "md: could not import %s!\n",
-			__bdevname(startdev, b));
+	if (IS_ERR(start_rdev))
 		return err;
-	}
+
 
 	/* NOTE: this can only work for 0.90.0 superblocks */
 	sb = (mdp_super_t*)page_address(start_rdev->sb_page);
@@ -1988,12 +1984,9 @@ static int autostart_array(dev_t startde
 		if (MAJOR(dev) != desc->major || MINOR(dev) != desc->minor)
 			continue;
 		rdev = md_import_device(dev, 0, 0);
-		if (IS_ERR(rdev)) {
-			printk(KERN_WARNING "md: could not import %s,"
-				" trying to run array nevertheless.\n",
-				__bdevname(dev, b));
+		if (IS_ERR(rdev))
 			continue;
-		}
+
 		list_add(&rdev->same_set, &pending_raid_disks);
 	}
 
@@ -2225,42 +2218,6 @@ static int add_new_disk(mddev_t * mddev,
 	return 0;
 }
 
-static int hot_generate_error(mddev_t * mddev, dev_t dev)
-{
-	char b[BDEVNAME_SIZE];
-	struct request_queue *q;
-	mdk_rdev_t *rdev;
-
-	if (!mddev->pers)
-		return -ENODEV;
-
-	printk(KERN_INFO "md: trying to generate %s error in %s ... \n",
-		__bdevname(dev, b), mdname(mddev));
-
-	rdev = find_rdev(mddev, dev);
-	if (!rdev) {
-		/* MD_BUG(); */ /* like hell - it's not a driver bug */
-		return -ENXIO;
-	}
-
-	if (rdev->desc_nr == -1) {
-		MD_BUG();
-		return -EINVAL;
-	}
-	if (!rdev->in_sync)
-		return -ENODEV;
-
-	q = bdev_get_queue(rdev->bdev);
-	if (!q) {
-		MD_BUG();
-		return -ENODEV;
-	}
-	printk(KERN_INFO "md: okay, generating error!\n");
-//	q->oneshot_error = 1; // disabled for now
-
-	return 0;
-}
-
 static int hot_remove_disk(mddev_t * mddev, dev_t dev)
 {
 	char b[BDEVNAME_SIZE];
@@ -2269,9 +2226,6 @@ static int hot_remove_disk(mddev_t * mdd
 	if (!mddev->pers)
 		return -ENODEV;
 
-	printk(KERN_INFO "md: trying to remove %s from %s ... \n",
-		__bdevname(dev, b), mdname(mddev));
-
 	rdev = find_rdev(mddev, dev);
 	if (!rdev)
 		return -ENXIO;
@@ -2299,9 +2253,6 @@ static int hot_add_disk(mddev_t * mddev,
 	if (!mddev->pers)
 		return -ENODEV;
 
-	printk(KERN_INFO "md: trying to hot-add %s to %s ... \n",
-		__bdevname(dev, b), mdname(mddev));
-
 	if (mddev->major_version != 0) {
 		printk(KERN_WARNING "%s: HOT_ADD may only be used with"
 			" version-0 superblocks.\n",
@@ -2561,7 +2512,6 @@ static int set_disk_faulty(mddev_t *mdde
 static int md_ioctl(struct inode *inode, struct file *file,
 			unsigned int cmd, unsigned long arg)
 {
-	char b[BDEVNAME_SIZE];
 	int err = 0;
 	void __user *argp = (void __user *)arg;
 	struct hd_geometry __user *loc = argp;
@@ -2620,8 +2570,7 @@ static int md_ioctl(struct inode *inode,
 		}
 		err = autostart_array(new_decode_dev(arg));
 		if (err) {
-			printk(KERN_WARNING "md: autostart %s failed!\n",
-				__bdevname(arg, b));
+			printk(KERN_WARNING "md: autostart failed!\n");
 			goto abort;
 		}
 		goto done;
@@ -2762,9 +2711,7 @@ static int md_ioctl(struct inode *inode,
 				err = add_new_disk(mddev, &info);
 			goto done_unlock;
 		}
-		case HOT_GENERATE_ERROR:
-			err = hot_generate_error(mddev, new_decode_dev(arg));
-			goto done_unlock;
+
 		case HOT_REMOVE_DISK:
 			err = hot_remove_disk(mddev, new_decode_dev(arg));
 			goto done_unlock;
@@ -3780,7 +3727,6 @@ void md_autodetect_dev(dev_t dev)
 
 static void autostart_arrays(int part)
 {
-	char b[BDEVNAME_SIZE];
 	mdk_rdev_t *rdev;
 	int i;
 
@@ -3790,11 +3736,9 @@ static void autostart_arrays(int part)
 		dev_t dev = detected_devices[i];
 
 		rdev = md_import_device(dev,0, 0);
-		if (IS_ERR(rdev)) {
-			printk(KERN_ALERT "md: could not import %s!\n",
-				__bdevname(dev, b));
+		if (IS_ERR(rdev))
 			continue;
-		}
+
 		if (rdev->faulty) {
 			MD_BUG();
 			continue;

^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 0 of 4] Introduction
@ 2004-09-03  2:20 NeilBrown
  0 siblings, 0 replies; 16+ messages in thread
From: NeilBrown @ 2004-09-03  2:20 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid

Following are 4 patches for md in 2.6.8.1-mm4

The first three are minor improvements and modifications either
required by or inspired by the fourth.

The fourth adds a new raid pers

^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 0 of 4] Introduction
@ 2004-11-02  3:37 NeilBrown
  0 siblings, 0 replies; 16+ messages in thread
From: NeilBrown @ 2004-11-02  3:37 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid

Following are 4 patches for md/raid against 2.6.10-rc1-mm2.

1/ Fix problem with linear arrays if component devices are > 2terabytes
2/ Fix data corruption in (experimental) RAID6 personality
3/ Fix possible oops with unplug_timer firing at the wrong time.
4/ Add new md personality "faulty".
    "Faulty" can be used to inject faults and so test failure modes
    of other raid levels and of filesystes.

NeilBrown


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH md 0 of 4] Introduction
@ 2005-03-08  5:50 NeilBrown
  2005-03-08  6:10 ` Andrew Morton
  2005-03-08 12:49 ` Peter T. Breuer
  0 siblings, 2 replies; 16+ messages in thread
From: NeilBrown @ 2005-03-08  5:50 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid


4 patches for md/raid in 2.6.11-mm1

The first two are trivial and should apply equally to 2.6.11

The second two fix bugs that were introduced by the recent 
bitmap-based-intent-logging patches and so are not relevant
to 2.6.11 yet. 

[PATCH md 1 of 4] Fix typo in super_1_sync
[PATCH md 2 of 4] Erroneous sizeof use in raid1
[PATCH md 3 of 4] Initialise sync_blocks in raid1 resync
[PATCH md 4 of 4] Fix md deadlock due to md thread processing delayed requests.

^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-08  5:50 NeilBrown
@ 2005-03-08  6:10 ` Andrew Morton
  2005-03-09  3:17   ` Neil Brown
  2005-03-08 12:49 ` Peter T. Breuer
  1 sibling, 1 reply; 16+ messages in thread
From: Andrew Morton @ 2005-03-08  6:10 UTC (permalink / raw)
  To: NeilBrown; +Cc: linux-raid

NeilBrown <neilb@cse.unsw.edu.au> wrote:
>
> The first two are trivial and should apply equally to 2.6.11
> 
>  The second two fix bugs that were introduced by the recent 
>  bitmap-based-intent-logging patches and so are not relevant
>  to 2.6.11 yet. 

The changelog for the "Fix typo in super_1_sync" patch doesn't actually say
what the patch does.  What are the user-visible consequences of not fixing
this?


Is the bitmap stuff now ready for Linus?

^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-08  5:50 NeilBrown
  2005-03-08  6:10 ` Andrew Morton
@ 2005-03-08 12:49 ` Peter T. Breuer
  2005-03-08 17:02   ` Paul Clements
  1 sibling, 1 reply; 16+ messages in thread
From: Peter T. Breuer @ 2005-03-08 12:49 UTC (permalink / raw)
  To: linux-raid

NeilBrown <neilb@cse.unsw.edu.au> wrote:
> The second two fix bugs that were introduced by the recent 
> bitmap-based-intent-logging patches and so are not relevant

Neil - can you describe for me (us all?) what is meant by
intent­logging here.

Well, I can guess - I suppose the driver marks the bitmap before a write
(or group of writes) and unmarks it when they have completed
successfully.  Is that it?

If so, how does it manage to mark what it is _going_ to do (without
psychic powers) on the disk bitmap?  Unmarking is easy - that needs a
queue of things due to be unmarked in the bitmap, and a point in time at
which they are all unmarked at once on disk.

Then resync would only deal with the marked blocks.

Peter

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-08 12:49 ` Peter T. Breuer
@ 2005-03-08 17:02   ` Paul Clements
  2005-03-08 19:05     ` Peter T. Breuer
  0 siblings, 1 reply; 16+ messages in thread
From: Paul Clements @ 2005-03-08 17:02 UTC (permalink / raw)
  To: Peter T. Breuer; +Cc: linux-raid

Peter T. Breuer wrote:

> Neil - can you describe for me (us all?) what is meant by
> intent­logging here.

Since I wrote a lot of the code, I guess I'll try...

> Well, I can guess - I suppose the driver marks the bitmap before a write
> (or group of writes) and unmarks it when they have completed
> successfully.  Is that it?

Yes. It marks the bitmap before writing (actually queues up the bitmap 
and normal writes in bunches for the sake of performance). The code is 
actually (loosely) based on your original bitmap (fr1) code.

> If so, how does it manage to mark what it is _going_ to do (without
> psychic powers) on the disk bitmap?  

That's actually fairly easy. The pages for the bitmap are locked in 
memory, so you just dirty the bits you want (which doesn't actually 
incur any I/O) and then when you're about to perform the normal writes, 
you flush the dirty bitmap pages to disk.

Once the writes are complete, a thread (we have the raid1d thread doing 
this) comes back along and flushes the (now clean) bitmap pages back to 
disk. If the pages get dirty again in the meantime (because of more 
I/O), we just leave them dirty and don't touch the disk.

> Then resync would only deal with the marked blocks.

Right. It clears the bitmap once things are back in sync.

--
Paul

-
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-08 17:02   ` Paul Clements
@ 2005-03-08 19:05     ` Peter T. Breuer
  2005-03-09  5:07       ` Neil Brown
  0 siblings, 1 reply; 16+ messages in thread
From: Peter T. Breuer @ 2005-03-08 19:05 UTC (permalink / raw)
  To: linux-raid

Paul Clements <paul.clements@steeleye.com> wrote:
> Peter T. Breuer wrote:
> > Neil - can you describe for me (us all?) what is meant by
> > intent-logging here.
> 
> Since I wrote a lot of the code, I guess I'll try...

Hi, Paul. Thanks.

> > Well, I can guess - I suppose the driver marks the bitmap before a write
> > (or group of writes) and unmarks it when they have completed
> > successfully.  Is that it?
> 
> Yes. It marks the bitmap before writing (actually queues up the bitmap 
> and normal writes in bunches for the sake of performance). The code is 
> actually (loosely) based on your original bitmap (fr1) code.

Yeah, I can see the traces.  I'm a little tired right now, but some
aspects of this idea vaguely worry me.  I'll see if I manage to
articulate those worries here despite my state. And you can dispell
them :).

Let me first of all guess at the intervals involved. I assume you will
write the marked parts of the bitmap to disk every 1/100th of a second or
so?  (I'd probably opt for 1/10th of a second or even every second just
to make sure it's not noticable on bandwidth and to heck with the
safety until we learn better what the tradeoffs are).  Or perhaps once
every hundred trasactions in busy times.

Now, there are races here.  You must mark the bitmap in memory before
every write, and unmark it after every complete write.  That is an
ordering constraint.  There is a race, however, to record the bitmap
state to disk.  Without any rendezvous or handshake or other
synchronization, one would simply be snapshotting the in-memory bitmap
to disk every so often, and the  on-disk bitmap would not always
accurately reflect the current state of completed transactions to the
mirror. The question is whether it shows an overly-pessimistic picture,
an overly-optimistic picture, or neither one nor the other.

I would naively imagine straight off that it cannot in general be
(appropriately) pessimistic because it does not know what writes will
occur in the next 1/100th second in order to be able to mark those on
the disk bitmap before they happen.  In the next section of your answer,
however, you say this is what happens, and therefore I deduce that

   a) 1/100th second's worth of writes to the mirror are first queued
   b) the in-memory bitmap is marked for these (if it exists as separate)
   c) the dirty parts of that bitmap are written to disk(s)
   d) the queued writes are carried out on the mirror
   e) the in-memory bitmap is unmarked for these
   f) the newly cleaned parts of that bitmap are written to disk. 

You may even have some sort of direct mapping between the on-disk
bitmap and the memory image, which could be quite effective, but
may run into problems with the address range available (bitmap must be
less than 2GB, no?), unless it maps only the necessary parts of the
bitmap at a time.  Well, if the kernel can manage that mapping window on
its own, it would be useful and probably what you have done.

But I digress. My immediate problem is that writes must be queued
first. I thought md traditionally did not queue requests, but instead
used its own make_request substitute to dispatch incoming requests as
they arrived.

Have you remodelled the md/raid1 make_request() fn?

And if so, do you also aggregate them? And what steps are taken to
preserve write ordering constraints (do some overlying file systems
still require these)?

> > If so, how does it manage to mark what it is _going_ to do (without
> > psychic powers) on the disk bitmap?  
> 
> That's actually fairly easy. The pages for the bitmap are locked in 
> memory,

That limits the size to about 2GB - oh, but perhaps you are doing as I
did and release bitmap pages when they are not dirty. Yes, you must.

>  so you just dirty the bits you want (which doesn't actually 
> incur any I/O) and then when you're about to perform the normal writes, 
> you flush the dirty bitmap pages to disk.

Hmm. I don't know how one can select pages to flush, but clearly one
can!  You maintain a list of dirtied pages, clearly. This list cannot be
larger than the list of outstanding requests. If you use the generic
kernel mechanisms, that will be 1000 or so, max.

> Once the writes are complete, a thread (we have the raid1d thread doing 
> this) comes back along and flushes the (now clean) bitmap pages back to 
> disk.

OK ..  there is a potential race here too, however, ...

> If the pages get dirty again in the meantime (because of more 
> I/O), we just leave them dirty and don't touch the disk.

Hmm. This appears to me to be an optimization. OK.

> > Then resync would only deal with the marked blocks.
> 
> Right. It clears the bitmap once things are back in sync.

Well, OK. Thinking it through as I write I see fewer problems. Thank
you for the explanation, and well done.

I have been meaning to merge the patches and see what comes out. I
presume you left out the mechanisms I included to allow a mirror
component to aggressively notify the array when it feels sick, and when
it feels better again. That required the array to be able to notify the
mirror components that they have been included in an array, and lodge
a callback hotline with  them.

Thanks again.

Peter


^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-08  6:10 ` Andrew Morton
@ 2005-03-09  3:17   ` Neil Brown
  2005-03-09  9:27     ` Mike Tran
  0 siblings, 1 reply; 16+ messages in thread
From: Neil Brown @ 2005-03-09  3:17 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-raid

On Monday March 7, akpm@osdl.org wrote:
> NeilBrown <neilb@cse.unsw.edu.au> wrote:
> >
> > The first two are trivial and should apply equally to 2.6.11
> > 
> >  The second two fix bugs that were introduced by the recent 
> >  bitmap-based-intent-logging patches and so are not relevant
> >  to 2.6.11 yet. 
> 
> The changelog for the "Fix typo in super_1_sync" patch doesn't actually say
> what the patch does.  What are the user-visible consequences of not fixing
> this?

-------
This fixes possible inconsistencies that might arise in a version-1 
superblock when devices fail and are removed.

Usage of version-1 superblocks is not yet widespread and no actual
problems have been reported.
--------
> 
> 
> Is the bitmap stuff now ready for Linus?

I agree with Paul - not yet.
I'd also like to get a bit more functionality in before it goes to
Linus, as the functionality may necessitate in interface change (I'm
not sure).
Specifically, I want the bitmap to be able to live near the superblock
rather than having to be in a file on a different filesystem.

NeilBrown

^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-08 19:05     ` Peter T. Breuer
@ 2005-03-09  5:07       ` Neil Brown
  2005-03-09 15:37         ` Peter T. Breuer
  0 siblings, 1 reply; 16+ messages in thread
From: Neil Brown @ 2005-03-09  5:07 UTC (permalink / raw)
  To: Peter T. Breuer; +Cc: linux-raid

On Tuesday March 8, ptb@lab.it.uc3m.es wrote:
> 
> But I digress. My immediate problem is that writes must be queued
> first. I thought md traditionally did not queue requests, but instead
> used its own make_request substitute to dispatch incoming requests as
> they arrived.
> 
> Have you remodelled the md/raid1 make_request() fn?

Somewhat.  Write requests are queued, and raid1d submits them when
it is happy that all bitmap updates have been done.

There is no '1/100th' second or anything like that.
When a write request arrives, the queue is 'plugged', requests are
queued, and bits in the in-memory bitmap are set.
When the queue is unplugged (by the filesystem or timeout) the bitmap
changes (if any) are flushed to disk, then the queued requests are
submitted. 

Bits on disk are cleaned lazily.


Note that for many applications, the bitmap does not need to be huge.
4K is enough for 1 bit per 2-3 megabytes on many large drives.
Having to sync 3 meg when just one block might be out-of-sync may seem
like a waste, but it is heaps better than syncing 100Gig!!

If a resync without bitmap logging takes 1 hour, I suspect a resync
with a 4K bitmap would have a good chance of finishing in under 1
minute (Depending on locality of references).  That is good enough for
me.

Of course, if one mirror is on the other side of the country, and a
normal sync requires 5 days over ADSL, then you would have a strong
case for a finer grained bitmap.

> 
> And if so, do you also aggregate them? And what steps are taken to
> preserve write ordering constraints (do some overlying file systems
> still require these)?

filesystems have never had any write ordering constraints, except that
IO must not be processed before it is requested, nor after it has been
acknowledged.  md continue to obey these restraints.

NeilBrown

^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-09  3:17   ` Neil Brown
@ 2005-03-09  9:27     ` Mike Tran
  0 siblings, 0 replies; 16+ messages in thread
From: Mike Tran @ 2005-03-09  9:27 UTC (permalink / raw)
  To: linux-raid

Hi Neil,

On Tue, 2005-03-08 at 21:17, Neil Brown wrote:
> On Monday March 7, akpm@osdl.org wrote:
> > NeilBrown <neilb@cse.unsw.edu.au> wrote:
> > >
> > > The first two are trivial and should apply equally to 2.6.11
> > > 
> > >  The second two fix bugs that were introduced by the recent 
> > >  bitmap-based-intent-logging patches and so are not relevant
> > >  to 2.6.11 yet. 
> > 
> > The changelog for the "Fix typo in super_1_sync" patch doesn't actually say
> > what the patch does.  What are the user-visible consequences of not fixing
> > this?
> 
> -------
> This fixes possible inconsistencies that might arise in a version-1 
> superblock when devices fail and are removed.
> 
> Usage of version-1 superblocks is not yet widespread and no actual
> problems have been reported.
> --------

EVMS 2.5.1 (http://evms.sf.net) has provided support for creation of MD
arrays using version-1 superblock.  Some of EVMS users actually tried to
use this new functionality.  You probably remember I posted a problem
and a patch to fix version-1 superblock update code.

We will continue to test and will report any problems.

--
Regards,
Mike T.



^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [PATCH md 0 of 4] Introduction
  2005-03-09  5:07       ` Neil Brown
@ 2005-03-09 15:37         ` Peter T. Breuer
  0 siblings, 0 replies; 16+ messages in thread
From: Peter T. Breuer @ 2005-03-09 15:37 UTC (permalink / raw)
  To: linux-raid

Neil Brown <neilb@cse.unsw.edu.au> wrote:
> On Tuesday March 8, ptb@lab.it.uc3m.es wrote:
> > Have you remodelled the md/raid1 make_request() fn?
> 
> Somewhat.  Write requests are queued, and raid1d submits them when
> it is happy that all bitmap updates have been done.

OK - so a slight modification of the kernel generic_make_request (I
haven't looked).  Mind you, I think that Paul said that just before
clearing bitmap entries, incoming requests were checked to see if a
bitmap entry should be marked again..

Perhaps both things happen. Bitmap pages in memory are updated as
clean after pending writes have finished and then marked as dirty as
necessary, and then flushed and when the flush finishes new accumulated
requests are started.

One can

> There is no '1/100th' second or anything like that.

I was trying in a way to give a definite image to what happens, rather
than speak abstractly. I'm sure that the ordinary kernel mechanism for
plugging and unplugging is used, as much as it is possible. If yu
unplug when the request struct reservoir is exhausted, then it will be
at 1K requests. If they are each 4KB, that will be every 4MB. At say
64MB/s, that will be every 1/16 s. And unplugging may happen more
frequently because of other kernel magic mumble mumble ...

> When a write request arrives, the queue is 'plugged', requests are
> queued, and bits in the in-memory bitmap are set.

OK.

> When the queue is unplugged (by the filesystem or timeout) the bitmap
> changes (if any) are flushed to disk, then the queued requests are
> submitted. 

That accumulates bitmap markings into the minimum number of extra
transactions.  It does impose extra latency, however.

I'm intrigued by exactly how you exert the memory pressure required to
force just the dirty bitmap pages out. I'll have to look it up.

> Bits on disk are cleaned lazily.

OK - so the disk bitmap state is always pessimistic. That's fine. Very
good.

> Note that for many applications, the bitmap does not need to be huge.
> 4K is enough for 1 bit per 2-3 megabytes on many large drives.
> Having to sync 3 meg when just one block might be out-of-sync may seem
> like a waste, but it is heaps better than syncing 100Gig!!

Yes - I used 1 bit per 1K, falling back to 1 bit per 2MB under memory
pressure.

> > And if so, do you also aggregate them? And what steps are taken to
> > preserve write ordering constraints (do some overlying file systems
> > still require these)?
> 
> filesystems have never had any write ordering constraints, except that
> IO must not be processed before it is requested, nor after it has been
> acknowledged.  md continue to obey these restraints.

Out of curiousity, is aggregation done on the queued requests? Or are
they all kept at 4KB? (or whatever - 1KB).

Thanks!

Peter


^ permalink raw reply	[flat|nested] 16+ messages in thread

end of thread, other threads:[~2005-03-09 15:37 UTC | newest]

Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2004-08-23  3:10 [PATCH md 0 of 4] Introduction NeilBrown
2004-08-23  3:10 ` [PATCH md 1 of 4] Assorted fixes/improvemnet to generic md resync code NeilBrown
2004-08-23  3:10 ` [PATCH md 4 of 4] RAID10 module for MD NeilBrown
2004-08-23  3:10 ` [PATCH md 3 of 4] Remove most calls to __bdevname from md.c NeilBrown
2004-08-23  3:10 ` [PATCH md 2 of 4] Assorted minor md/raid1 fixes NeilBrown
  -- strict thread matches above, loose matches on Subject: below --
2004-09-03  2:20 [PATCH md 0 of 4] Introduction NeilBrown
2004-11-02  3:37 NeilBrown
2005-03-08  5:50 NeilBrown
2005-03-08  6:10 ` Andrew Morton
2005-03-09  3:17   ` Neil Brown
2005-03-09  9:27     ` Mike Tran
2005-03-08 12:49 ` Peter T. Breuer
2005-03-08 17:02   ` Paul Clements
2005-03-08 19:05     ` Peter T. Breuer
2005-03-09  5:07       ` Neil Brown
2005-03-09 15:37         ` Peter T. Breuer

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox