* Help RAID5 reshape Oops / backup-file
@ 2007-10-09 18:58 Nagilum
0 siblings, 0 replies; 11+ messages in thread
From: Nagilum @ 2007-10-09 18:58 UTC (permalink / raw)
To: linux-raid
[-- Attachment #1.1: Type: text/plain, Size: 9923 bytes --]
Hi,
During the process of reshaping a Raid5 from 3 (/dev/sd[a-c]) to 5
devices (/dev/sd[a-e]) the system was accidentally shut down.
I know I was stupid I should have used a --backup-file but stupid me didn't.
Thanks for not rubbing it any further. :(
Ok, here is what I have:
nas:~# uname -a
Linux nas 2.6.18-5-amd64 #1 SMP Thu Aug 30 01:14:54 UTC 2007 x86_64 GNU/Linux
nas:~# mdadm --version
mdadm - v2.5.6 - 9 November 2006
nas:~# mdadm -Q --detail /dev/md0
/dev/md0:
Version : 00.91.03
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Persistence : Superblock is persistent
Update Time : Mon Oct 8 23:59:27 2007
State : active, degraded, Not Started
Active Devices : 3
Working Devices : 5
Failed Devices : 0
Spare Devices : 2
Layout : left-symmetric
Chunk Size : 16K
Delta Devices : 2, (3->5)
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Events : 0.470134
Number Major Minor RaidDevice State
0 8 0 0 active sync /dev/sda
1 8 16 1 active sync /dev/sdb
2 8 32 2 active sync /dev/sdc
3 0 0 3 removed
4 0 0 4 removed
5 8 48 - spare /dev/sdd
6 8 64 - spare /dev/sde
nas:~# mdadm -E /dev/sd[a-e]
/dev/sda:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f425054d - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 0 8 0 0 active sync /dev/sda
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sdb:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f425055f - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 1 8 16 1 active sync /dev/sdb
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sdc:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f4250571 - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 2 8 32 2 active sync /dev/sdc
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sdd:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f42505b9 - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 5 8 48 -1 spare /dev/sdd
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sde:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f42505db - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 6 8 64 -1 spare /dev/sde
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
nas:~# mdadm /dev/md0 -r /dev/sde
mdadm: hot remove failed for /dev/sde: No such device
nas:~# cat /proc/mdstat
Personalities : [raid6] [raid5] [raid4]
md0 : inactive sda[0] sdd[5](S) sde[6](S) sdc[2] sdb[1]
2441543360 blocks super 0.91
unused devices: <none>
So reshaping was almost done.
The way I imagine how reshaping works you'd basically have an already
remapped area growing from the start of the drives (which grows during
remapping) and some not-yet-remapped-area which spans from the end of
the drives towards wherever remapping is currently reading data from.
At the beginning of the processing it equals the start of the drives
but it shrinks faster than the remapped area grows. So there is a
growing gap between the two areas which contains still original
unmapped data.
The point is, as soon as this area grows large enough the backup-file
should become unneeded.
Ok, now if mdadm wants me to provide that file I should also be able
to re-create it using the "Reshape pos'n" and the drive geometry (and
dd).
Now the question is how to do it?
I also have build mdadm-2.6.3 which appears to see things more clearly:
nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 /dev/sd[a-e]
mdadm: Failed to restore critical section for reshape, sorry.
So if I could create the backup file I should be able to continue..
Any help would be greatly appreciated!
Alexander.
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #1.2.1: Type: multipart/signed, Size: 0 bytes --]
[-- Attachment #1.2.2: Type: text/plain, Size: 9663 bytes --]
Hi,
During the process of reshaping a Raid5 from 3 (/dev/sd[a-c]) to 5
devices (/dev/sd[a-e]) the system was accidentally shut down.
I know I was stupid I should have used a --backup-file but stupid me didn't.
Thanks for not rubbing it any further. :(
Ok, here is what I have:
nas:~# uname -a
Linux nas 2.6.18-5-amd64 #1 SMP Thu Aug 30 01:14:54 UTC 2007 x86_64 GNU/Linux
nas:~# mdadm --version
mdadm - v2.5.6 - 9 November 2006
nas:~# mdadm -Q --detail /dev/md0
/dev/md0:
Version : 00.91.03
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Persistence : Superblock is persistent
Update Time : Mon Oct 8 23:59:27 2007
State : active, degraded, Not Started
Active Devices : 3
Working Devices : 5
Failed Devices : 0
Spare Devices : 2
Layout : left-symmetric
Chunk Size : 16K
Delta Devices : 2, (3->5)
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Events : 0.470134
Number Major Minor RaidDevice State
0 8 0 0 active sync /dev/sda
1 8 16 1 active sync /dev/sdb
2 8 32 2 active sync /dev/sdc
3 0 0 3 removed
4 0 0 4 removed
5 8 48 - spare /dev/sdd
6 8 64 - spare /dev/sde
nas:~# mdadm -E /dev/sd[a-e]
/dev/sda:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f425054d - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 0 8 0 0 active sync /dev/sda
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sdb:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f425055f - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 1 8 16 1 active sync /dev/sdb
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sdc:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f4250571 - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 2 8 32 2 active sync /dev/sdc
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sdd:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f42505b9 - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 5 8 48 -1 spare /dev/sdd
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
/dev/sde:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f42505db - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 6 8 64 -1 spare /dev/sde
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
nas:~# mdadm /dev/md0 -r /dev/sde
mdadm: hot remove failed for /dev/sde: No such device
nas:~# cat /proc/mdstat
Personalities : [raid6] [raid5] [raid4]
md0 : inactive sda[0] sdd[5](S) sde[6](S) sdc[2] sdb[1]
2441543360 blocks super 0.91
unused devices: <none>
So reshaping was almost done.
The way I imagine how reshaping works you'd basically have an already
remapped area growing from the start of the drives (which grows during
remapping) and some not-yet-remapped-area which spans from the end of
the drives towards wherever remapping is currently reading data from.
At the beginning of the processing it equals the start of the drives
but it shrinks faster than the remapped area grows. So there is a
growing gap between the two areas which contains still original
unmapped data.
The point is, as soon as this area grows large enough the backup-file
should become unneeded.
Ok, now if mdadm wants me to provide that file I should also be able
to re-create it using the "Reshape pos'n" and the drive geometry (and
dd).
Now the question is how to do it?
I also have build mdadm-2.6.3 which appears to see things more clearly:
nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 /dev/sd[a-e]
mdadm: Failed to restore critical section for reshape, sorry.
So if I could create the backup file I should be able to continue..
Any help would be greatly appreciated!
Alexander.
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #1.2.3: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
@ 2007-10-11 12:25 Nagilum
2007-10-11 23:51 ` Neil Brown
0 siblings, 1 reply; 11+ messages in thread
From: Nagilum @ 2007-10-11 12:25 UTC (permalink / raw)
To: linux-raid; +Cc: Neil Brown
[-- Attachment #1: Type: text/plain, Size: 10832 bytes --]
Ok, after looking in "Grow.c" I can see that the backup file is
removed once the critial section has passed:
if (backup_file)
unlink(backup_file);
printf(Name ": ... critical section passed.\n");
Since I had passed that point I'll try to find out where
Grow_restart() stumbles. By looking at it I'm not even sure it's able
to "resume" and not just restart. :-/
----- Message from nagilum@nagilum.org ---------
Date: Tue, 09 Oct 2007 20:58:47 +0200
From: Nagilum <nagilum@nagilum.org>
Reply-To: Nagilum <nagilum@nagilum.org>
Subject: Help RAID5 reshape Oops / backup-file
To: linux-raid@vger.kernel.org
> Hi,
> During the process of reshaping a Raid5 from 3 (/dev/sd[a-c]) to 5
> devices (/dev/sd[a-e]) the system was accidentally shut down.
> I know I was stupid I should have used a --backup-file but stupid me didn't.
> Thanks for not rubbing it any further. :(
> Ok, here is what I have:
>
> nas:~# uname -a
> Linux nas 2.6.18-5-amd64 #1 SMP Thu Aug 30 01:14:54 UTC 2007 x86_64 GNU/Linux
> nas:~# mdadm --version
> mdadm - v2.5.6 - 9 November 2006
> nas:~# mdadm -Q --detail /dev/md0
> /dev/md0:
> Version : 00.91.03
> Creation Time : Sat Sep 15 21:11:41 2007
> Raid Level : raid5
> Device Size : 488308672 (465.69 GiB 500.03 GB)
> Raid Devices : 5
> Total Devices : 5
> Preferred Minor : 0
> Persistence : Superblock is persistent
>
> Update Time : Mon Oct 8 23:59:27 2007
> State : active, degraded, Not Started
> Active Devices : 3
> Working Devices : 5
> Failed Devices : 0
> Spare Devices : 2
>
> Layout : left-symmetric
> Chunk Size : 16K
>
> Delta Devices : 2, (3->5)
>
> UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
> Events : 0.470134
>
> Number Major Minor RaidDevice State
> 0 8 0 0 active sync /dev/sda
> 1 8 16 1 active sync /dev/sdb
> 2 8 32 2 active sync /dev/sdc
> 3 0 0 3 removed
> 4 0 0 4 removed
>
> 5 8 48 - spare /dev/sdd
> 6 8 64 - spare /dev/sde
>
>
> nas:~# mdadm -E /dev/sd[a-e]
> /dev/sda:
> Magic : a92b4efc
> Version : 00.91.00
> UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
> Creation Time : Sat Sep 15 21:11:41 2007
> Raid Level : raid5
> Device Size : 488308672 (465.69 GiB 500.03 GB)
> Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
> Raid Devices : 5
> Total Devices : 5
> Preferred Minor : 0
>
> Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
> Delta Devices : 2 (3->5)
>
> Update Time : Mon Oct 8 23:59:27 2007
> State : clean
> Active Devices : 5
> Working Devices : 5
> Failed Devices : 0
> Spare Devices : 0
> Checksum : f425054d - correct
> Events : 0.470134
>
> Layout : left-symmetric
> Chunk Size : 16K
>
> Number Major Minor RaidDevice State
> this 0 8 0 0 active sync /dev/sda
>
> 0 0 8 0 0 active sync /dev/sda
> 1 1 8 16 1 active sync /dev/sdb
> 2 2 8 32 2 active sync /dev/sdc
> 3 3 8 64 3 active sync /dev/sde
> 4 4 8 48 4 active sync /dev/sdd
> /dev/sdb:
> Magic : a92b4efc
> Version : 00.91.00
> UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
> Creation Time : Sat Sep 15 21:11:41 2007
> Raid Level : raid5
> Device Size : 488308672 (465.69 GiB 500.03 GB)
> Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
> Raid Devices : 5
> Total Devices : 5
> Preferred Minor : 0
>
> Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
> Delta Devices : 2 (3->5)
>
> Update Time : Mon Oct 8 23:59:27 2007
> State : clean
> Active Devices : 5
> Working Devices : 5
> Failed Devices : 0
> Spare Devices : 0
> Checksum : f425055f - correct
> Events : 0.470134
>
> Layout : left-symmetric
> Chunk Size : 16K
>
> Number Major Minor RaidDevice State
> this 1 8 16 1 active sync /dev/sdb
>
> 0 0 8 0 0 active sync /dev/sda
> 1 1 8 16 1 active sync /dev/sdb
> 2 2 8 32 2 active sync /dev/sdc
> 3 3 8 64 3 active sync /dev/sde
> 4 4 8 48 4 active sync /dev/sdd
> /dev/sdc:
> Magic : a92b4efc
> Version : 00.91.00
> UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
> Creation Time : Sat Sep 15 21:11:41 2007
> Raid Level : raid5
> Device Size : 488308672 (465.69 GiB 500.03 GB)
> Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
> Raid Devices : 5
> Total Devices : 5
> Preferred Minor : 0
>
> Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
> Delta Devices : 2 (3->5)
>
> Update Time : Mon Oct 8 23:59:27 2007
> State : clean
> Active Devices : 5
> Working Devices : 5
> Failed Devices : 0
> Spare Devices : 0
> Checksum : f4250571 - correct
> Events : 0.470134
>
> Layout : left-symmetric
> Chunk Size : 16K
>
> Number Major Minor RaidDevice State
> this 2 8 32 2 active sync /dev/sdc
>
> 0 0 8 0 0 active sync /dev/sda
> 1 1 8 16 1 active sync /dev/sdb
> 2 2 8 32 2 active sync /dev/sdc
> 3 3 8 64 3 active sync /dev/sde
> 4 4 8 48 4 active sync /dev/sdd
> /dev/sdd:
> Magic : a92b4efc
> Version : 00.91.00
> UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
> Creation Time : Sat Sep 15 21:11:41 2007
> Raid Level : raid5
> Device Size : 488308672 (465.69 GiB 500.03 GB)
> Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
> Raid Devices : 5
> Total Devices : 5
> Preferred Minor : 0
>
> Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
> Delta Devices : 2 (3->5)
>
> Update Time : Mon Oct 8 23:59:27 2007
> State : clean
> Active Devices : 5
> Working Devices : 5
> Failed Devices : 0
> Spare Devices : 0
> Checksum : f42505b9 - correct
> Events : 0.470134
>
> Layout : left-symmetric
> Chunk Size : 16K
>
> Number Major Minor RaidDevice State
> this 5 8 48 -1 spare /dev/sdd
>
> 0 0 8 0 0 active sync /dev/sda
> 1 1 8 16 1 active sync /dev/sdb
> 2 2 8 32 2 active sync /dev/sdc
> 3 3 8 64 3 active sync /dev/sde
> 4 4 8 48 4 active sync /dev/sdd
> /dev/sde:
> Magic : a92b4efc
> Version : 00.91.00
> UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
> Creation Time : Sat Sep 15 21:11:41 2007
> Raid Level : raid5
> Device Size : 488308672 (465.69 GiB 500.03 GB)
> Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
> Raid Devices : 5
> Total Devices : 5
> Preferred Minor : 0
>
> Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
> Delta Devices : 2 (3->5)
>
> Update Time : Mon Oct 8 23:59:27 2007
> State : clean
> Active Devices : 5
> Working Devices : 5
> Failed Devices : 0
> Spare Devices : 0
> Checksum : f42505db - correct
> Events : 0.470134
>
> Layout : left-symmetric
> Chunk Size : 16K
>
> Number Major Minor RaidDevice State
> this 6 8 64 -1 spare /dev/sde
>
> 0 0 8 0 0 active sync /dev/sda
> 1 1 8 16 1 active sync /dev/sdb
> 2 2 8 32 2 active sync /dev/sdc
> 3 3 8 64 3 active sync /dev/sde
> 4 4 8 48 4 active sync /dev/sdd
>
> nas:~# mdadm /dev/md0 -r /dev/sde
> mdadm: hot remove failed for /dev/sde: No such device
> nas:~# cat /proc/mdstat
> Personalities : [raid6] [raid5] [raid4]
> md0 : inactive sda[0] sdd[5](S) sde[6](S) sdc[2] sdb[1]
> 2441543360 blocks super 0.91
>
> unused devices: <none>
>
> So reshaping was almost done.
> The way I imagine how reshaping works you'd basically have an already
> remapped area growing from the start of the drives (which grows during
> remapping) and some not-yet-remapped-area which spans from the end of
> the drives towards wherever remapping is currently reading data from.
> At the beginning of the processing it equals the start of the drives
> but it shrinks faster than the remapped area grows. So there is a
> growing gap between the two areas which contains still original
> unmapped data.
> The point is, as soon as this area grows large enough the backup-file
> should become unneeded.
> Ok, now if mdadm wants me to provide that file I should also be able
> to re-create it using the "Reshape pos'n" and the drive geometry (and
> dd).
> Now the question is how to do it?
> I also have build mdadm-2.6.3 which appears to see things more clearly:
>
> nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 /dev/sd[a-e]
> mdadm: Failed to restore critical section for reshape, sorry.
>
> So if I could create the backup file I should be able to continue..
> Any help would be greatly appreciated!
> Alexander.
----- End message from nagilum@nagilum.org -----
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-11 12:25 Help RAID5 reshape Oops / backup-file Nagilum
@ 2007-10-11 23:51 ` Neil Brown
2007-10-12 6:43 ` Nagilum
0 siblings, 1 reply; 11+ messages in thread
From: Neil Brown @ 2007-10-11 23:51 UTC (permalink / raw)
To: Nagilum; +Cc: linux-raid
On Thursday October 11, nagilum@nagilum.org wrote:
> Ok, after looking in "Grow.c" I can see that the backup file is
> removed once the critial section has passed:
>
> if (backup_file)
> unlink(backup_file);
>
> printf(Name ": ... critical section passed.\n");
>
> Since I had passed that point I'll try to find out where
> Grow_restart() stumbles. By looking at it I'm not even sure it's able
> to "resume" and not just restart. :-/
>
It isn't a problem that you didn't specify a backup-file.
If you don't, mdadm uses some spare space on one of the new drives.
After the critical section has passed, the backup file isn't needed
any longer.
The problem is that mdadm still wants to find and recover from it.
I throughly tested mdadm restarting from a crash during the critical
section, but it looks like I didn't properly test restarting from a
later crash.
I think if you just change the 'return 1' at the end of Grow_restart
to 'return 0' it should work for you.
I'll try to get this fixed properly (and tested) and release a 2.6.4.
NeilBrown
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-11 23:51 ` Neil Brown
@ 2007-10-12 6:43 ` Nagilum
2007-10-14 16:55 ` Nagilum
0 siblings, 1 reply; 11+ messages in thread
From: Nagilum @ 2007-10-12 6:43 UTC (permalink / raw)
To: Neil Brown; +Cc: linux-raid
[-- Attachment #1: Type: text/plain, Size: 2815 bytes --]
----- Message from neilb@suse.de ---------
Date: Fri, 12 Oct 2007 09:51:08 +1000
From: Neil Brown <neilb@suse.de>
Reply-To: Neil Brown <neilb@suse.de>
Subject: Re: Help RAID5 reshape Oops / backup-file
To: Nagilum <nagilum@nagilum.org>
Cc: linux-raid@vger.kernel.org
> On Thursday October 11, nagilum@nagilum.org wrote:
>> Ok, after looking in "Grow.c" I can see that the backup file is
>> removed once the critial section has passed:
>>
>> if (backup_file)
>> unlink(backup_file);
>>
>> printf(Name ": ... critical section passed.\n");
>>
>> Since I had passed that point I'll try to find out where
>> Grow_restart() stumbles. By looking at it I'm not even sure it's able
>> to "resume" and not just restart. :-/
>>
>
> It isn't a problem that you didn't specify a backup-file.
> If you don't, mdadm uses some spare space on one of the new drives.
> After the critical section has passed, the backup file isn't needed
> any longer.
> The problem is that mdadm still wants to find and recover from it.
>
> I throughly tested mdadm restarting from a crash during the critical
> section, but it looks like I didn't properly test restarting from a
> later crash.
>
> I think if you just change the 'return 1' at the end of Grow_restart
> to 'return 0' it should work for you.
>
> I'll try to get this fixed properly (and tested) and release a 2.6.4.
>
> NeilBrown
>
----- End message from neilb@suse.de -----
Thanks, I changed Grow_restart as suggested, now I get:
nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 /dev/sd[a-e]
mdadm: /dev/md0 assembled from 3 drives and 2 spares - not enough to
start the array.
nas:~/mdadm-2.6.3# cat /proc/mdstat
Personalities : [raid6] [raid5] [raid4]
md0 : inactive sda[0] sde[6] sdd[5] sdc[2] sdb[1]
2441543360 blocks
unused devices: <none>
which is similar to what the old mdadm is telling me.
I'll try to find out where it gets the idea these are spares..
Would it be a good idea to update to vanilla 2.6.23 instead of running
Debian Etch's 2.6.18-5?
If there is anything I can do to help with v2.6.4 let me know!
Thanks,
Alex.
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-12 6:43 ` Nagilum
@ 2007-10-14 16:55 ` Nagilum
2007-10-14 23:31 ` Neil Brown
0 siblings, 1 reply; 11+ messages in thread
From: Nagilum @ 2007-10-14 16:55 UTC (permalink / raw)
To: linux-raid; +Cc: Neil Brown
[-- Attachment #1: Type: text/plain, Size: 5251 bytes --]
Can someone tell me if I'm on the right track?
I've now noticed the following:
# ~/mdadm-2.6.3/mdadm -v -A /dev/md0 /dev/sd[d-e]
mdadm: looking for devices for /dev/md0
mdadm: /dev/sdd is identified as a member of /dev/md0, slot -1.
mdadm: /dev/sde is identified as a member of /dev/md0, slot -1.
mdadm: No suitable drives found for /dev/md0
This "slot -1" is also visible in the examine output:
# mdadm -E /dev/sdd
/dev/sdd:
Magic : a92b4efc
Version : 00.91.00
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Device Size : 488308672 (465.69 GiB 500.03 GB)
Array Size : 1953234688 (1862.75 GiB 2000.11 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Reshape pos'n : 872095808 (831.70 GiB 893.03 GB)
Delta Devices : 2 (3->5)
Update Time : Mon Oct 8 23:59:27 2007
State : clean
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Checksum : f42505b9 - correct
Events : 0.470134
Layout : left-symmetric
Chunk Size : 16K
Number Major Minor RaidDevice State
this 5 8 48 -1 spare /dev/sdd
0 0 8 0 0 active sync /dev/sda
1 1 8 16 1 active sync /dev/sdb
2 2 8 32 2 active sync /dev/sdc
3 3 8 64 3 active sync /dev/sde
4 4 8 48 4 active sync /dev/sdd
So can someone confirm that is the likely source of my problem?
And hopefully - if that is indeed the problem - someone can tell me
how to update that slot number?
Thanks,
Alex.
----- Message from nagilum@nagilum.org ---------
Date: Fri, 12 Oct 2007 08:43:28 +0200
From: Nagilum <nagilum@nagilum.org>
Reply-To: Nagilum <nagilum@nagilum.org>
Subject: Re: Help RAID5 reshape Oops / backup-file
To: Neil Brown <neilb@suse.de>
Cc: linux-raid@vger.kernel.org
> ----- Message from neilb@suse.de ---------
> Date: Fri, 12 Oct 2007 09:51:08 +1000
> From: Neil Brown <neilb@suse.de>
> Reply-To: Neil Brown <neilb@suse.de>
> Subject: Re: Help RAID5 reshape Oops / backup-file
> To: Nagilum <nagilum@nagilum.org>
> Cc: linux-raid@vger.kernel.org
>
>
>> It isn't a problem that you didn't specify a backup-file.
>> If you don't, mdadm uses some spare space on one of the new drives.
>> After the critical section has passed, the backup file isn't needed
>> any longer.
>> The problem is that mdadm still wants to find and recover from it.
>>
>> I throughly tested mdadm restarting from a crash during the critical
>> section, but it looks like I didn't properly test restarting from a
>> later crash.
>>
>> I think if you just change the 'return 1' at the end of Grow_restart
>> to 'return 0' it should work for you.
>>
>> I'll try to get this fixed properly (and tested) and release a 2.6.4.
>>
>> NeilBrown
>>
>
>
> ----- End message from neilb@suse.de -----
>
> Thanks, I changed Grow_restart as suggested, now I get:
> nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 /dev/sd[a-e]
> mdadm: /dev/md0 assembled from 3 drives and 2 spares - not enough to
> start the array.
>
> nas:~/mdadm-2.6.3# cat /proc/mdstat
> Personalities : [raid6] [raid5] [raid4]
> md0 : inactive sda[0] sde[6] sdd[5] sdc[2] sdb[1]
> 2441543360 blocks
>
> unused devices: <none>
>
> which is similar to what the old mdadm is telling me.
> I'll try to find out where it gets the idea these are spares..
> Would it be a good idea to update to vanilla 2.6.23 instead of running
> Debian Etch's 2.6.18-5?
> If there is anything I can do to help with v2.6.4 let me know!
> Thanks,
> Alex.
>
> ========================================================================
> # _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
> # / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
> # / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
> # /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
> # /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
> ========================================================================
>
>
> ----------------------------------------------------------------
> cakebox.homeunix.net - all the machine one needs..
----- End message from nagilum@nagilum.org -----
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-14 16:55 ` Nagilum
@ 2007-10-14 23:31 ` Neil Brown
2007-10-15 11:55 ` Nagilum
0 siblings, 1 reply; 11+ messages in thread
From: Neil Brown @ 2007-10-14 23:31 UTC (permalink / raw)
To: Nagilum; +Cc: linux-raid
On Sunday October 14, nagilum@nagilum.org wrote:
> Can someone tell me if I'm on the right track?
> I've now noticed the following:
> # ~/mdadm-2.6.3/mdadm -v -A /dev/md0 /dev/sd[d-e]
> mdadm: looking for devices for /dev/md0
> mdadm: /dev/sdd is identified as a member of /dev/md0, slot -1.
> mdadm: /dev/sde is identified as a member of /dev/md0, slot -1.
> mdadm: No suitable drives found for /dev/md0
Hmm... that might be useful..
I just found your earlier email where you said:
> After the machine came back up (on a rescue disk) I thought I'd
> simply have to go through the process again. So I use add add the
> new disk again.
> Although that worked, I am now unable to resume the growing
> process.
Using "add add" again was not correct, and should not have been
possible.
You should have simply assembled the array with the full new set of
devices. Then reshape would have automatically restarted properly.
Can you remember *exactly* what you did? If I can reproduce the
situation, I can find the best way to fix it and send you something to
try.
NeilBrown
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-14 23:31 ` Neil Brown
@ 2007-10-15 11:55 ` Nagilum
2007-10-15 12:08 ` Nagilum
2007-10-16 1:16 ` Neil Brown
0 siblings, 2 replies; 11+ messages in thread
From: Nagilum @ 2007-10-15 11:55 UTC (permalink / raw)
To: Neil Brown; +Cc: linux-raid
[-- Attachment #1: Type: text/plain, Size: 4600 bytes --]
----- Message from neilb@suse.de ---------
Date: Mon, 15 Oct 2007 09:31:23 +1000
From: Neil Brown <neilb@suse.de>
Reply-To: Neil Brown <neilb@suse.de>
Subject: Re: Help RAID5 reshape Oops / backup-file
To: Nagilum <nagilum@nagilum.org>
Cc: linux-raid@vger.kernel.org
> On Sunday October 14, nagilum@nagilum.org wrote:
>> Can someone tell me if I'm on the right track?
>> I've now noticed the following:
>> # ~/mdadm-2.6.3/mdadm -v -A /dev/md0 /dev/sd[d-e]
>> mdadm: looking for devices for /dev/md0
>> mdadm: /dev/sdd is identified as a member of /dev/md0, slot -1.
>> mdadm: /dev/sde is identified as a member of /dev/md0, slot -1.
>> mdadm: No suitable drives found for /dev/md0
>
> Hmm... that might be useful..
>
> I just found your earlier email where you said:
>
>> After the machine came back up (on a rescue disk) I thought I'd
>> simply have to go through the process again. So I use add add the
>> new disk again.
>> Although that worked, I am now unable to resume the growing
>> process.
>
> Using "add add" again was not correct, and should not have been
> possible.
> You should have simply assembled the array with the full new set of
> devices. Then reshape would have automatically restarted properly.
>
> Can you remember *exactly* what you did? If I can reproduce the
> situation, I can find the best way to fix it and send you something to
> try.
>
> NeilBrown
>
----- End message from neilb@suse.de -----
Sure, here it goes:
The system is running Debian Etch ia64, kernel 2.6.18,
(since the exact versions might be important in this case I made
copies of what I deemed to be relevant available online)
a copy of the "linux/drivers/md" folder of that particular kernel can
be found at:
http://www.nagilum.de/md/md
Etch comes with mdadm-2.5.6 + Debian patches.
See http://www.nagilum.de/md/mdadm-2.5.6/debian/changelog
I made the whole Debian Package available here:
http://www.nagilum.de/md/
- "mdadm-2.5.6" the extracted source with Debian patches applied
- mdadm_2.5.6-9.diff.gz the diff to mdadm_2.5.6.orig.tar.gz
- mdadm_2.5.6-9_i386.deb the i385 version of the package, however I
was/am using mdadm_2.5.6-9_ia64.deb
- "mdadm_2.5.6-9.dsc" description file for building the .deb
The Raid was being reshaped from three to five drives when the
shutdown was issued. I assume the shutdown went normally since the
machine was off and there was no power interruption.
Upon booting the system it became apparent that the RAID was non functional.
The system boots off of a USB stick and then mounts its root
filesystem from the RAID. Assembling the RAID happens within the
initrd. The relevant scripts can be found here:
http://www.nagilum.de/md/local-top/
I booted a rescue disk which is based on the identical Linux version.
I looked at the "mdadm -Q --detail /dev/md0" output and saw only 3 of
the 5 disks in the RAID. Then I did (what I should not have done) the
add of the two new disks, assuming that mdadm will touch these in a
harmful way (without using --force) and refuse to do so if that's not
the way to add active disk.
The disks were added but the reshape did not continue.
Up until now I can't think of anything else I did that could have
changed something. (and "mdadm -Q --detail /dev/md0" looks the same
ever since)
I think, what I should have done instead of adding those disks would
have been to either use --re-add and/or update /etc/mdadm/mdadm.conf.
But then again I never expected this to become so problematic. :(
By now I can also boot with 2.6.23 (I'll update to 2.6.23.1 shortly)
and I have the latest mdadm tools (in parallel to the old ones).
I also build the test_stripe utility and tried a very briefly the
"test" argument, but it wanted me to specify an existing file so I
chickened out. ;)
Thanks a lot for looking into this!
Alex.
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-15 11:55 ` Nagilum
@ 2007-10-15 12:08 ` Nagilum
2007-10-16 1:16 ` Neil Brown
1 sibling, 0 replies; 11+ messages in thread
From: Nagilum @ 2007-10-15 12:08 UTC (permalink / raw)
To: linux-raid; +Cc: Neil Brown
[-- Attachment #1: Type: text/plain, Size: 1239 bytes --]
----- Message from nagilum@nagilum.org ---------
Date: Mon, 15 Oct 2007 13:55:22 +0200
From: Nagilum <nagilum@nagilum.org>
Reply-To: Nagilum <nagilum@nagilum.org>
Subject: Re: Help RAID5 reshape Oops / backup-file
To: Neil Brown <neilb@suse.de>
Cc: linux-raid@vger.kernel.org
> - mdadm_2.5.6-9_i386.deb the i385 version of the package, however I
i386 of course
> add of the two new disks, assuming that mdadm will touch these in a
will _not_ touch these in a harmful way
----- End message from nagilum@nagilum.org -----
..stupid typos ;)
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-15 11:55 ` Nagilum
2007-10-15 12:08 ` Nagilum
@ 2007-10-16 1:16 ` Neil Brown
2007-10-16 12:50 ` Nagilum
1 sibling, 1 reply; 11+ messages in thread
From: Neil Brown @ 2007-10-16 1:16 UTC (permalink / raw)
To: Nagilum; +Cc: linux-raid
Thanks for the extra details.
I still cannot manage to reproduce it which is frustrating, but I
think I can fix your array for you.
Get the source for mdadm 2.6.3, apply the following patch, then use
mdadm -A /dev/md0 --update=this /dev/sd[abcde]
that should re-write the part of the superblocks that is wrong, then
assemble the array.
Please let me know how it goes.
Also, if you could show me "mdadm.conf" and "mdrun.conf" from the
initrd, that might help.
Thanks,
NeilBrown
diff --git a/Grow.c b/Grow.c
index 825747e..8ad1537 100644
--- a/Grow.c
+++ b/Grow.c
@@ -978,5 +978,5 @@ int Grow_restart(struct supertype *st, struct mdinfo *info, int *fdlist, int cnt
/* And we are done! */
return 0;
}
- return 1;
+ return 0;
}
diff --git a/mdadm.c b/mdadm.c
index 40fdccf..7e7e803 100644
--- a/mdadm.c
+++ b/mdadm.c
@@ -584,6 +584,8 @@ int main(int argc, char *argv[])
exit(2);
}
update = optarg;
+ if (strcmp(update, "this")==0)
+ continue;
if (strcmp(update, "sparc2.2")==0)
continue;
if (strcmp(update, "super-minor") == 0)
diff --git a/super0.c b/super0.c
index 0396c2c..e33e623 100644
--- a/super0.c
+++ b/super0.c
@@ -394,6 +394,21 @@ static int update_super0(struct mdinfo *info, void *sbv, char *update,
fprintf (stderr, Name ": adjusting superblock of %s for 2.2/sparc compatability.\n",
devname);
}
+ if (strcmp(update, "this") == 0) {
+ /* to fix a particular corrupt superblock.
+ */
+ int i;
+ for (i=0; i<10; i++)
+ if (sb->disks[i].major == sb->this_disk.major &&
+ sb->disks[i].minor == sb->this_disk.minor) {
+ if (sb->this_disk.number == sb->disks[i].number)
+ break;
+ fprintf(stderr, Name ": Setting this disk from %d to %d\n",
+ sb->this_disk.number, sb->disks[i].number);
+ sb->this_disk = sb->disks[i];
+ break;
+ }
+ }
if (strcmp(update, "super-minor") ==0) {
sb->md_minor = info->array.md_minor;
if (verbose > 0)
^ permalink raw reply related [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-16 1:16 ` Neil Brown
@ 2007-10-16 12:50 ` Nagilum
2007-10-17 13:13 ` Nagilum
0 siblings, 1 reply; 11+ messages in thread
From: Nagilum @ 2007-10-16 12:50 UTC (permalink / raw)
To: Neil Brown; +Cc: linux-raid
[-- Attachment #1: Type: text/plain, Size: 4985 bytes --]
----- Message from neilb@suse.de ---------
Date: Tue, 16 Oct 2007 11:16:19 +1000
From: Neil Brown <neilb@suse.de>
Reply-To: Neil Brown <neilb@suse.de>
Subject: Re: Help RAID5 reshape Oops / backup-file
To: Nagilum <nagilum@nagilum.org>
Cc: linux-raid@vger.kernel.org
>
> Thanks for the extra details.
> I still cannot manage to reproduce it which is frustrating, but I
> think I can fix your array for you.
>
> Get the source for mdadm 2.6.3, apply the following patch, then use
>
> mdadm -A /dev/md0 --update=this /dev/sd[abcde]
>
> that should re-write the part of the superblocks that is wrong, then
> assemble the array.
>
> Please let me know how it goes.
>
> Also, if you could show me "mdadm.conf" and "mdrun.conf" from the
> initrd, that might help.
>
> Thanks,
> NeilBrown
>
----- End message from neilb@suse.de -----
Thanks a bunch mate!
So far it looks very well:
nas:~/mdadm-2.6.3# ./mdadm -Q --detail /dev/md0
/dev/md0:
Version : 00.91.03
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Used Dev Size : 488308672 (465.69 GiB 500.03 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Persistence : Superblock is persistent
Update Time : Mon Oct 8 23:59:27 2007
State : active, degraded, Not Started
Active Devices : 3
Working Devices : 5
Failed Devices : 0
Spare Devices : 2
Layout : left-symmetric
Chunk Size : 16K
Delta Devices : 2, (3->5)
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Events : 0.470134
Number Major Minor RaidDevice State
0 8 0 0 active sync /dev/sda
1 8 16 1 active sync /dev/sdb
2 8 32 2 active sync /dev/sdc
3 0 0 3 removed
4 0 0 4 removed
5 8 48 - spare /dev/sdd
6 8 64 - spare /dev/sde
nas:~/mdadm-2.6.3# mdadm -S /dev/md0
mdadm: stopped /dev/md0
nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 --update=this /dev/sd[abcde]
mdadm: Setting this disk from 5 to 4
mdadm: Setting this disk from 6 to 3
mdadm: /dev/md0 assembled from 3 drives and 2 spares - not enough to
start the array.
nas:~/mdadm-2.6.3# mdadm -S /dev/md0
mdadm: stopped /dev/md0
nas:~/mdadm-2.6.3# ./mdadm -A /dev/md0 /dev/sd[abcde]
mdadm: /dev/md0 has been started with 5 drives.
nas:~/mdadm-2.6.3# ./mdadm -Q --detail /dev/md0
/dev/md0:
Version : 00.91.03
Creation Time : Sat Sep 15 21:11:41 2007
Raid Level : raid5
Array Size : 976617344 (931.37 GiB 1000.06 GB)
Used Dev Size : 488308672 (465.69 GiB 500.03 GB)
Raid Devices : 5
Total Devices : 5
Preferred Minor : 0
Persistence : Superblock is persistent
Update Time : Tue Oct 16 13:42:03 2007
State : clean, recovering
Active Devices : 5
Working Devices : 5
Failed Devices : 0
Spare Devices : 0
Layout : left-symmetric
Chunk Size : 16K
Reshape Status : 44% complete
Delta Devices : 2, (3->5)
UUID : 25da80a6:d56eb9d6:0d7656f3:2f233380
Events : 0.470212
Number Major Minor RaidDevice State
0 8 0 0 active sync /dev/sda
1 8 16 1 active sync /dev/sdb
2 8 32 2 active sync /dev/sdc
3 8 64 3 active sync /dev/sde
4 8 48 4 active sync /dev/sdd
nas:~# cat /proc/mdstat
md0 : active raid5 sda[0] sdd[4] sde[3] sdc[2] sdb[1]
976617344 blocks super 0.91 level 5, 16k chunk, algorithm 2
[5/5] [UUUUU]
[=============>.......] reshape = 67.5% (329927392/488308672)
finish=48.6min speed=54235K/sec
unused devices: <none>
I'll send an update (including the configs) when its done and I've
verified everything is healthy. (and my heart has stopped racing ;)
At first I was bit scared because of the order (sde before sdd) but
that's consistent with the mdadm -E output from the devices earlier,
so it looks like I'll soon have my data back. *yay* :)
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: Help RAID5 reshape Oops / backup-file
2007-10-16 12:50 ` Nagilum
@ 2007-10-17 13:13 ` Nagilum
0 siblings, 0 replies; 11+ messages in thread
From: Nagilum @ 2007-10-17 13:13 UTC (permalink / raw)
To: Neil Brown; +Cc: linux-raid
[-- Attachment #1: Type: text/plain, Size: 2028 bytes --]
----- Message from nagilum@nagilum.org ---------
Date: Tue, 16 Oct 2007 14:50:09 +0200
From: Nagilum <nagilum@nagilum.org>
Reply-To: Nagilum <nagilum@nagilum.org>
Subject: Re: Help RAID5 reshape Oops / backup-file
To: Neil Brown <neilb@suse.de>
Cc: linux-raid@vger.kernel.org
> ----- Message from neilb@suse.de ---------
> Date: Tue, 16 Oct 2007 11:16:19 +1000
> From: Neil Brown <neilb@suse.de>
> Reply-To: Neil Brown <neilb@suse.de>
> Subject: Re: Help RAID5 reshape Oops / backup-file
> To: Nagilum <nagilum@nagilum.org>
> Cc: linux-raid@vger.kernel.org
>
>> Please let me know how it goes.
>>
>> Also, if you could show me "mdadm.conf" and "mdrun.conf" from the
>> initrd, that might help.
>>
>> Thanks,
>> NeilBrown
>>
>
>
> ----- End message from neilb@suse.de -----
Ok, the array reshaped successfully and is back in production. :)
Here the content of the (old) initrd /etc/mdadm/mdadm.conf:
DEVICE partitions
ARRAY /dev/md0 level=raid5 num-devices=3
UUID=25da80a6:d56eb9d6:c7780c0e:bc15422d
That needed updating of course. I don't have a "mdrun.conf" on my
system or on the initrd.
I'll try to replicate the issue using a different machine and plain
files (lets see if that works) at the weekend and let you know if I
succeed.
Again, thank you so much for the patch!
Alex.
========================================================================
# _ __ _ __ http://www.nagilum.org/ \n icq://69646724 #
# / |/ /__ ____ _(_) /_ ____ _ nagilum@nagilum.org \n +491776461165 #
# / / _ `/ _ `/ / / // / ' \ Amiga (68k/PPC): AOS/NetBSD/Linux #
# /_/|_/\_,_/\_, /_/_/\_,_/_/_/_/ Mac (PPC): MacOS-X / NetBSD /Linux #
# /___/ x86: FreeBSD/Linux/Solaris/Win2k ARM9: EPOC EV6 #
========================================================================
----------------------------------------------------------------
cakebox.homeunix.net - all the machine one needs..
[-- Attachment #2: PGP Digital Signature --]
[-- Type: application/pgp-signature, Size: 187 bytes --]
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2007-10-17 13:13 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2007-10-11 12:25 Help RAID5 reshape Oops / backup-file Nagilum
2007-10-11 23:51 ` Neil Brown
2007-10-12 6:43 ` Nagilum
2007-10-14 16:55 ` Nagilum
2007-10-14 23:31 ` Neil Brown
2007-10-15 11:55 ` Nagilum
2007-10-15 12:08 ` Nagilum
2007-10-16 1:16 ` Neil Brown
2007-10-16 12:50 ` Nagilum
2007-10-17 13:13 ` Nagilum
-- strict thread matches above, loose matches on Subject: below --
2007-10-09 18:58 Nagilum
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox