* Re: All drive in Raid 5 are in 'spare' mode
From: Phil Turmel @ 2015-02-16 20:03 UTC (permalink / raw)
To: Dush; +Cc: linux-raid@vger.kernel.org
In-Reply-To: <CAL7hTOfhWiCmKVVCtpmHuB-OJkZjQsuaKZMN8oFKAArh=NuSHQ@mail.gmail.com>
Hi Dush,
{Convention on kernel.org is to trim replies and either bottom post or
interleave. Please don't top-post.}
On 02/16/2015 02:38 PM, Dush wrote:
> Hi Phil,
>
> Thanks for your answer!
>
> Unfortunately, I think I just loosed a disk (sde)... I don't see it
> anymore in /dev , I have in dmesg:
> d
> [ 12.280021] ata7: softreset failed (1st FIS failed)
> [ 22.280019] ata7: softreset failed (1st FIS failed)
> [ 57.280015] ata7: softreset failed (1st FIS failed)
> [ 57.280222] ata7: limiting SATA link speed to 1.5 Gbps
> [ 62.453345] ata7: softreset failed (device not ready)
> [ 62.453558] ata7: reset failed, giving up
Where are the forensics I asked for as "Step one"? Did you read about
and fix any timeout mismatch issue?
[trim /]
> You was right, I already tried to start the raid and it succeed to do
> it with 3 drives: b, c and e. Then I added the d because I thought it
> was de-synchronized.
> Now I think my drive e was out of this raid for a while and I started
> to had trouble because d started to had some issues.
>
> Is it possible to force raid to start with b, c and d (forcing d to be
> 'normal')? Time for me to copy everything to another drive...
No. sdd3 was converted to a spare.
Also, your device names have changed. You *must* keep track of which
one is which "RaidDevice". Specifically, what was "sde" now appears to
be "sdd". Did you reboot? You know that device names are not
guaranteed to be consistent from one boot to the next, I hope. Show an
excerpt from "ls -l /dev/disk/by-id/" with your next report so we know
which drive serial number has which name.
Phil
^ permalink raw reply
* Re: What are mdadm maintainers to do?
From: Phil Turmel @ 2015-02-16 19:44 UTC (permalink / raw)
To: Chris, linux-raid
In-Reply-To: <loom.20150216T183419-387@post.gmane.org>
On 02/16/2015 12:48 PM, Chris wrote:
>
> Thank you for the additional information, it calls for action.
>
>
> OK, calling for a solution to stop desktop drives from causing data loss and
> affecting the mdadm reputation:
>
>
> I gather that mdadm could ship with one additional udev rule that calls a
> script to check/set scterc, or falls back to increasing the system timout.
>
> Phil, you mentioned having posted such a script, could you prepare it for
> addition to the mdadm package?
No, I've posted snippets for users to customize in their own rc.local or
distro equivalent. I vaguely recall posting a generic script for some
common cases, but I've personally converted to raid-rated drives
everywhere in the past couple years.
Somebody else will have to tackle this.
> Would maintainers be ok with adding such a udev rule and script to the package?
Not my call, but keep in mind that this will add a dependency on
smartmontools or whatever means is used to access/write to scterc.
Phil
^ permalink raw reply
* Re: All drive in Raid 5 are in 'spare' mode
From: Dush @ 2015-02-16 19:38 UTC (permalink / raw)
To: Phil Turmel; +Cc: linux-raid@vger.kernel.org
In-Reply-To: <54DC04C1.4010107@turmel.org>
Hi Phil,
Thanks for your answer!
Unfortunately, I think I just loosed a disk (sde)... I don't see it
anymore in /dev , I have in dmesg:
d
[ 12.280021] ata7: softreset failed (1st FIS failed)
[ 22.280019] ata7: softreset failed (1st FIS failed)
[ 57.280015] ata7: softreset failed (1st FIS failed)
[ 57.280222] ata7: limiting SATA link speed to 1.5 Gbps
[ 62.453345] ata7: softreset failed (device not ready)
[ 62.453558] ata7: reset failed, giving up
And my reports look like this now:
# cat /proc/mdstat
Personalities : [raid6] [raid5] [raid4]
md126 : inactive sdb3[3](S) sdd3[1](S) sdc3[0](S)
1447416000 blocks
md127 : active (auto-read-only) raid5 sdd2[1] sdb2[3] sdc2[0]
16530624 blocks level 5, 64k chunk, algorithm 2 [4/3] [UU_U]
unused devices: <none>
# mdadm --examine /dev/sd[b-e]3
/dev/sdb3:
Magic : a92b4efc
Version : 0.90.00
UUID : 3327f442:a00b59b2:1397f3c2:236c0edf
Creation Time : Tue Jan 27 13:03:52 2009
Raid Level : raid5
Used Dev Size : 482472000 (460.12 GiB 494.05 GB)
Array Size : 1447416000 (1380.36 GiB 1482.15 GB)
Raid Devices : 4
Total Devices : 4
Preferred Minor : 126
Update Time : Wed Jan 21 20:55:48 2015
State : active
Active Devices : 3
Working Devices : 4
Failed Devices : 1
Spare Devices : 1
Checksum : 6e656c69 - correct
Events : 49656
Layout : left-symmetric
Chunk Size : 64K
Number Major Minor RaidDevice State
this 3 8 19 3 active sync /dev/sdb3
0 0 8 35 0 active sync /dev/sdc3
1 1 8 67 1 active sync
2 2 0 0 2 faulty removed
3 3 8 19 3 active sync /dev/sdb3
4 4 8 51 4 spare /dev/sdd3
/dev/sdc3:
Magic : a92b4efc
Version : 0.90.00
UUID : 3327f442:a00b59b2:1397f3c2:236c0edf
Creation Time : Tue Jan 27 13:03:52 2009
Raid Level : raid5
Used Dev Size : 482472000 (460.12 GiB 494.05 GB)
Array Size : 1447416000 (1380.36 GiB 1482.15 GB)
Raid Devices : 4
Total Devices : 4
Preferred Minor : 126
Update Time : Wed Jan 21 23:34:52 2015
State : clean
Active Devices : 2
Working Devices : 3
Failed Devices : 2
Spare Devices : 1
Checksum : 6e6653d5 - correct
Events : 49666
Layout : left-symmetric
Chunk Size : 64K
Number Major Minor RaidDevice State
this 0 8 35 0 active sync /dev/sdc3
0 0 8 35 0 active sync /dev/sdc3
1 1 8 67 1 active sync
2 2 0 0 2 faulty removed
3 3 0 0 3 faulty removed
4 4 8 51 4 spare /dev/sdd3
/dev/sdd3:
Magic : a92b4efc
Version : 0.90.00
UUID : 3327f442:a00b59b2:1397f3c2:236c0edf
Creation Time : Tue Jan 27 13:03:52 2009
Raid Level : raid5
Used Dev Size : 482472000 (460.12 GiB 494.05 GB)
Array Size : 1447416000 (1380.36 GiB 1482.15 GB)
Raid Devices : 4
Total Devices : 4
Preferred Minor : 126
Update Time : Wed Jan 21 23:34:52 2015
State : clean
Active Devices : 2
Working Devices : 3
Failed Devices : 2
Spare Devices : 1
Checksum : 6e6653f7 - correct
Events : 49666
Layout : left-symmetric
Chunk Size : 64K
Number Major Minor RaidDevice State
this 1 8 67 1 active sync
0 0 8 35 0 active sync /dev/sdc3
1 1 8 67 1 active sync
2 2 0 0 2 faulty removed
3 3 0 0 3 faulty removed
4 4 8 51 4 spare /dev/sdd3
You was right, I already tried to start the raid and it succeed to do
it with 3 drives: b, c and e. Then I added the d because I thought it
was de-synchronized.
Now I think my drive e was out of this raid for a while and I started
to had trouble because d started to had some issues.
Is it possible to force raid to start with b, c and d (forcing d to be
'normal')? Time for me to copy everything to another drive...
Thanks,
Dush
On 12 February 2015 at 01:41, Phil Turmel <philip@turmel.org> wrote:
> Hi Dush,
>
> On 02/11/2015 02:56 PM, Dush wrote:
>> Hi,
>>
>> I have a RAID 5 composed by 4x 500Go hdd but for some days, it's 'inactive'.
>>
>> I'm not raid expert and I prefer asking before doing an unrecoverable mistake...
>>
>> Is it possible to fix this raid (md126)?
>> Is it possible to recover data on it?
>
> Probably. Very good report, btw.
>
>> Do I have a disk to change or it's "just" a desynchronization between disks?
>
> One disk is now truly a spare (/dev/sdd3), which suggests you already
> tried to '--add' it and didn't get anywhere.
>
> Step one: collect some forensics for later. syslog or dmesg containing
> your failure events. Can be trimmed to just device and md stuff.
> "smartctl -x /dev/sdX" for each drive involved in the arrays.
>
> Then, we'll try the simple stuff.
>
> Make sure the array is stopped with:
>
> mdadm --stop /dev/md126
>
> Then, force assemble it without sdd:
>
> mdadm --assemble --force --verbose --run /dev/md126 /dev/sd[bce]3
>
> If that works, mount it and catch a backup of critical files.
>
> Then add your /dev/sdd3 back to the array and let it rebuild:
>
> mdadm --add /dev/md126 /dev/sdd3
>
> It may not make it through the rebuild if you have the common timeout
> mismatch problem.[1] Show the dmesg and smartctl data (pasted inline is
> preferred) and we'll see.
>
> Phil
>
> Recent typical case:
> [1] http://marc.info/?l=linux-raid&m=142353387024935&w=1
>
^ permalink raw reply
* What are mdadm maintainers to do? (was: desktop disk's error recovery timeouts)
From: Chris @ 2015-02-16 17:48 UTC (permalink / raw)
To: linux-raid
In-Reply-To: <54E226B5.1080500@turmel.org>
Thank you for the additional information, it calls for action.
OK, calling for a solution to stop desktop drives from causing data loss and
affecting the mdadm reputation:
I gather that mdadm could ship with one additional udev rule that calls a
script to check/set scterc, or falls back to increasing the system timout.
Phil, you mentioned having posted such a script, could you prepare it for
addition to the mdadm package?
Would maintainers be ok with adding such a udev rule and script to the package?
Kind Regards,
Chris
^ permalink raw reply
* Re: desktop disk's error recovery timouts
From: Phil Turmel @ 2015-02-16 17:19 UTC (permalink / raw)
To: Chris, linux-raid
In-Reply-To: <loom.20150216T165111-551@post.gmane.org>
On 02/16/2015 11:15 AM, Chris wrote:
> Phil, thank you for dropping in with this hint. It very likly applies to
> the disks in the docking station. I searched the mailing list, most hits
> said to search for the keywords, though. ;-)
I don't always have time to explain. :-(
> To understand the issue, I think
> https://en.wikipedia.org/wiki/Error_recovery_control
> was good.
Good starting points in the archives:
http://marc.info/?l=linux-raid&m=135811522817345&w=1
http://marc.info/?l=linux-raid&m=133761065622164&w=2
http://marc.info/?l=linux-raid&m=135863964624202&w=2
http://marc.info/?l=linux-raid&m=139050322510249&w=2
There's useful info in each entire thread, though.
Phil
^ permalink raw reply
* desktop disk's error recovery timouts (was: re-add POLICY)
From: Chris @ 2015-02-16 16:15 UTC (permalink / raw)
To: linux-raid
In-Reply-To: <54E1EDEA.1030503@turmel.org>
Phil Turmel <philip <at> turmel.org> writes:
> On 02/16/2015 07:23 AM, Chris wrote:
> > .... with raid members that got pulled and are save to
> > re-sync. (e.g. after the occasional bad block error that gets remapped by
> > the hardrives firmware)
>
> This should not be part of your concern here, as MD will handle
> occassional UREs by reconstructing them and rewriting them on the fly,
Phil, thank you for dropping in with this hint. It very likly applies to
the disks in the docking station. I searched the mailing list, most hits
said to search for the keywords, though. ;-)
To understand the issue, I think
https://en.wikipedia.org/wiki/Error_recovery_control
was good.
It would be good if this configuration information could be available there
or at https://raid.wiki.kernel.org
Cheers,
Chris
----
I compiled some snippets from your messages, that could serve as a basis to
correction/completion by someone knowledgeable:
The default linux controller timeout is 30 seconds. Drives
that spend longer than the timeout in recovery will be reset. If they
don't respond to the reset (because they're busy in recovery) when the
raid tries to write the correct data back to them, they will be kicked
out of the array.
You *must* set ERC shorter than the
timeout, or set the driver timeout longer than the drive's worst-case
recovery time. The defaults for desktop drives are *not* suitable for
linux software raid.
I strongly encourage you to run "smartctl -l scterc /dev/sdX" for each
of your drives. For any drive that warns that it doesn't support SCT
ERC, set the controller device timeout to 180 like so:
echo 180 >/sys/block/sdX/device/timeout
If the report says read or write ERC is disabled, run "smartctl -l
scterc,70,70 /dev/sdX" to set it to 7.0 seconds.
You then set up a boot-time script to do these adjustments at every restart,
and make sure you performing regular scrub runs to ...?
You might not want that kind of long device timeout, but then you shouldn't
use desktop drives in md RAID.
Anyone using desktop drives which don't support SCT ERC in md RAID is
liable to see long timeouts on the simplest bad sector, and they
probably prefer to keep the drive in the array AND have the sector
rewritten after reconstruction than have the drive failed out of the array.
^ permalink raw reply
* Re: re-add POLICY
From: Phil Turmel @ 2015-02-16 13:17 UTC (permalink / raw)
To: Chris, linux-raid
In-Reply-To: <loom.20150216T124230-883@post.gmane.org>
Hi Chris,
On 02/16/2015 07:23 AM, Chris wrote:
> .... with raid members that got pulled and are save to
> re-sync. (e.g. after the occasional bad block error that gets remapped by
> the hardrives firmware)
This should not be part of your concern here, as MD will handle
occassional UREs by reconstructing them and rewriting them on the fly,
-- without failing the device. If devices are failing after read
errors, you have a different problem. (Hint: look at recent threads
for "timeout mismatch".)
Phil
^ permalink raw reply
* Re: RAID 1 metadata - keep separate from mirror disks ?
From: Phil Turmel @ 2015-02-16 13:02 UTC (permalink / raw)
To: Suresh Babu Kandukuru, linux-raid; +Cc: Ankur Bose
In-Reply-To: <17135927-75bb-499a-8f43-6748c872feb0@default>
Hi Suresh,
On 02/16/2015 06:37 AM, Suresh Babu Kandukuru wrote:
> Yeh . Thank Phil . it is quite useful . Now we see there are two
> options . 1) without metadata 2) externally managed metadata . But we
> would like to keep the metadata external to mirror leg ( like on host
> local drive ) for the reasons mentioned below , not the externally
> managed metadata , just by keeping metadata on mirror legs . Do we
> have any option with md driver ?
I'm not sure I understand your follow-up question. You've restated my
answer, and re-iterated your preference to not have metadata on the
member devices.
So you need to write a service to handle metadata the way you want. Or
write scripts that will do --build operations at appropriate times
without metadata.
I'm not entirely clear why the on-member metadata is unacceptable, but
as such, either remaining option needs some code of your own.
Phil
^ permalink raw reply
* Re: re-add POLICY
From: Chris @ 2015-02-16 12:23 UTC (permalink / raw)
To: linux-raid
In-Reply-To: <20150216142845.0d50207c@notabene.brown>
NeilBrown <neilb <at> suse.de> writes:
> Does your array have a write-intent bitmap configured?
> If it does, then "POLICY action=re-add" really should work.
Thank you for your insight. You are correct, the array has no write-intent
bitmap.
> If it doesn't, then maybe you need "POLICY action=spare".
OK, I will test this when the notebook is back in the house.
Actually, the man page had kind of kept me from trying this, because it
mentions the condition "if the device is bare", and I didn't want arbitrary
bare disk, partition, or free space to be automatically added, but just to
trigger an automatic try with raid members that got pulled and are save to
re-sync. (e.g. after the occasional bad block error that gets remapped by
the hardrives firmware)
[man page: spare works] "as above and additionally: if the device is
bare it can become a spare if there is any array that it is a candidate for
based on domains and metadata."
Also, I wouldn't want a temporarily removed raid member to be added as spare
to some other array. Only have them added (re-synced even if no bitmap
re-add is possible) to the array they belong according to their superblock.
> This isn't the default, because depending on exactly how/why the device
> failed, it may not be safe to treat it as a spare.
OK, I can imagine detecting the corner cases may require some inteligent
error logging.
What I am looking for is a safe re-sync configuration option between
bitmap-based re-add, and treating a device as arbitrary spare drive.
Practically, this could be something like an additional action=re-sync
option in between re-add/spare, or having the "re-add" action also do
(non-bitmap) full re-syncs, if the device is in a clean state.
May recording the fail event count in the remaining superblocks help, as
described in
http://permalink.gmane.org/gmane.linux.raid/48077
help to detect the clean state?
Kind Regards,
Chris
^ permalink raw reply
* RE: RAID 1 metadata - keep separate from mirror disks ?
From: Suresh Babu Kandukuru @ 2015-02-16 11:37 UTC (permalink / raw)
To: Phil Turmel, linux-raid; +Cc: Ankur Bose
In-Reply-To: <54DB6BF0.3040309@turmel.org>
Yeh . Thank Phil . it is quite useful . Now we see there are two options . 1) without metadata 2) externally managed metadata . But we would like to keep the metadata external to mirror leg ( like on host local drive ) for the reasons mentioned below , not the externally managed metadata , just by keeping metadata on mirror legs
. Do we have any option with md driver ?
/Suresh
Principal developer | Hyderabad Team Lead
Phone: +914067246370 | Mobile: +919701451727
Oracle Maxrep development
ORACLE India Hyderabad
Oracle is committed to developing practices and products that help protect the environment
-----Original Message-----
From: Phil Turmel [mailto:philip@turmel.org]
Sent: Wednesday, February 11, 2015 8:19 PM
To: Suresh Babu Kandukuru; linux-raid@vger.kernel.org
Subject: Re: RAID 1 metadata - keep separate from mirror disks ?
Good morning Suresh,
On 02/11/2015 07:14 AM, Suresh Babu Kandukuru wrote:
> Hi There,
>
> On the RAID 1 metadata: is there any way to keep the metadata
> separate from the mirror disks? Could you guide us on this ?,
> please. In general, we need to keep all metadata off the device
> itself, leaving all the device available for user data. This is
> particularly important in the migration case, where we want to take an
> existing LUN and add a second leg to it to create the mirror device
> without changing any of the data or metadata on the LUN.
If you look at "man 4 md" you'll see some options. If a legacy array type meets your needs, you can operate without metadata at all. Use "mdadm --build" to assemble your raid at each boot.
Or, if your storage server can insert a leg ahead of you current LUN, you can then create the array with an explicit data offset matching the size of the inserted leg. Create it degraded with the existing LUN, then add (a) LUN(s) to start mirroring. This process will leave you the option to resize with more legs later.
Or you can add a leg to the end and create your array with version 1.0 metadata, which is placed at the end of the device.
Finally, you could write your own metadata container service for use with mdmon. (That's a bit beyond my ability, sorry.)
Phil
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html
^ permalink raw reply
* Re: mdadm failed to remove internal bitmap
From: NeilBrown @ 2015-02-16 6:48 UTC (permalink / raw)
To: gary; +Cc: linux-raid
In-Reply-To: <54E1914C.9030000@gmail.com>
[-- Attachment #1: Type: text/plain, Size: 1127 bytes --]
On Mon, 16 Feb 2015 14:42:20 +0800 gary <gary.mdjiang@gmail.com> wrote:
> Hi Neil,
>
> Please check the followings.
> > What kernel are you running?
> linux48:~ # uname -r
> 3.12.32-33-default
Does this have any patches on top of 3.12.32 that touch md.c ?
>
> And 3.12.28-4-default kernel is ok.
There are no differences between 3.12.28 and 3.12.32 that could affect this.
> ioctl(3, RAID_VERSION, 0x7fff2e10c060) = 0
> ioctl(3, GET_BITMAP_FILE, 0x7fff2e10c1f0) = 0
> ioctl(3, GET_ARRAY_INFO, 0x7fff2e10c1a0) = 0
> ioctl(3, SET_ARRAY_INFO, 0x7fff2e10c1a0) = -1 EINVAL (Invalid argument)
> write(2, "mdadm: failed to remove internal"..., 41mdadm: failed to
> remove internal bitmap.
EINVAL from SET_ARRAY_INFO almost certainly comes from update_array_info().
It can happen if:
- more than 1 thing needs to be updated - seems unlikely
- pers->quiesce is NULL - not possible for raid1.
- mddev->bitmap->storage.file is not NULL. Seems unlikely.
I suggest you look at the code you are actually running, and possible add
some printks to tell you where it is failing.
NeilBrown
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 811 bytes --]
^ permalink raw reply
* Re: mdadm failed to remove internal bitmap
From: gary @ 2015-02-16 6:42 UTC (permalink / raw)
To: NeilBrown; +Cc: linux-raid
In-Reply-To: <20150216143419.553cf93e@notabene.brown>
Hi Neil,
Please check the followings.
> What kernel are you running?
linux48:~ # uname -r
3.12.32-33-default
And 3.12.28-4-default kernel is ok.
> Please use "strace" on mdadm in a case where it fails, and post the result.
linux48:~ # strace mdadm --grow --bitmap=none /dev/md127
execve("/sbin/mdadm", ["mdadm", "--grow", "--bitmap=none",
"/dev/md127"], [/* 58 vars */]) = 0
brk(0) = 0xfe8000
mmap(NULL, 4096, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0)
= 0x7fcab07ad000
access("/etc/ld.so.preload", R_OK) = -1 ENOENT (No such file or
directory)
open("/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
fstat(3, {st_mode=S_IFREG|0644, st_size=93919, ...}) = 0
mmap(NULL, 93919, PROT_READ, MAP_PRIVATE, 3, 0) = 0x7fcab0796000
close(3) = 0
open("/lib64/libc.so.6", O_RDONLY|O_CLOEXEC) = 3
read(3,
"\177ELF\2\1\1\0\0\0\0\0\0\0\0\0\3\0>\0\1\0\0\0\20\34\2\0\0\0\0\0"...,
832) = 832
fstat(3, {st_mode=S_IFREG|0755, st_size=1978611, ...}) = 0
mmap(NULL, 3832352, PROT_READ|PROT_EXEC, MAP_PRIVATE|MAP_DENYWRITE, 3,
0) = 0x7fcab01e6000
mprotect(0x7fcab0384000, 2097152, PROT_NONE) = 0
mmap(0x7fcab0584000, 24576, PROT_READ|PROT_WRITE,
MAP_PRIVATE|MAP_FIXED|MAP_DENYWRITE, 3, 0x19e000) = 0x7fcab0584000
mmap(0x7fcab058a000, 14880, PROT_READ|PROT_WRITE,
MAP_PRIVATE|MAP_FIXED|MAP_ANONYMOUS, -1, 0) = 0x7fcab058a000
close(3) = 0
mmap(NULL, 4096, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0)
= 0x7fcab0795000
mmap(NULL, 4096, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0)
= 0x7fcab0794000
mmap(NULL, 4096, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0)
= 0x7fcab0793000
arch_prctl(ARCH_SET_FS, 0x7fcab0794700) = 0
mprotect(0x7fcab0584000, 16384, PROT_READ) = 0
mprotect(0x677000, 4096, PROT_READ) = 0
mprotect(0x7fcab07ae000, 4096, PROT_READ) = 0
munmap(0x7fcab0796000, 93919) = 0
getpid() = 6000
brk(0) = 0xfe8000
brk(0x1009000) = 0x1009000
open("/dev/md127", O_RDWR) = 3
fstat(3, {st_mode=S_IFBLK|0660, st_rdev=makedev(9, 127), ...}) = 0
ioctl(3, RAID_VERSION, 0x7fff2e10d160) = 0
open("/etc/mdadm.conf", O_RDONLY) = -1 ENOENT (No such file or
directory)
open("/etc/mdadm/mdadm.conf", O_RDONLY) = -1 ENOENT (No such file or
directory)
open("/etc/mdadm.conf.d", O_RDONLY) = -1 ENOENT (No such file or
directory)
uname({sys="Linux", node="linux48", ...}) = 0
geteuid() = 0
fstat(3, {st_mode=S_IFBLK|0660, st_rdev=makedev(9, 127), ...}) = 0
ioctl(3, RAID_VERSION, 0x7fff2e10c060) = 0
ioctl(3, GET_BITMAP_FILE, 0x7fff2e10c1f0) = 0
ioctl(3, GET_ARRAY_INFO, 0x7fff2e10c1a0) = 0
ioctl(3, SET_ARRAY_INFO, 0x7fff2e10c1a0) = -1 EINVAL (Invalid argument)
write(2, "mdadm: failed to remove internal"..., 41mdadm: failed to
remove internal bitmap.
) = 41
exit_group(1) = ?
+++ exited with 1 +++
Thanks,
gary
^ permalink raw reply
* Re: [md PATCH] md/raid1: round up to bdev_logical_block_size in narrow_write_error
From: NeilBrown @ 2015-02-16 3:54 UTC (permalink / raw)
To: Nate Dailey; +Cc: linux-raid
In-Reply-To: <54DCDC91.6010808@stratus.com>
[-- Attachment #1: Type: text/plain, Size: 1668 bytes --]
On Thu, 12 Feb 2015 12:02:09 -0500 Nate Dailey <nate.dailey@stratus.com>
wrote:
> This modifies raid1's narrow_write_error to round up block_sectors to the
> device's logical block size.
>
> This prevents sd complaining about "Bad block number requested" for non-512-byte
> sector disks.
>
> Signed-off-by: Nate Dailey <nate.dailey@stratus.com>
> ---
>
> diff -Nupr a/drivers/md/raid1.c b/drivers/md/raid1.c
> --- a/drivers/md/raid1.c 2015-02-10 15:29:02.000000000 -0500
> +++ b/drivers/md/raid1.c 2015-02-10 15:29:45.000000000 -0500
> @@ -2206,7 +2206,8 @@ static int narrow_write_error(struct r1b
> if (rdev->badblocks.shift < 0)
> return 0;
>
> - block_sectors = 1 << rdev->badblocks.shift;
> + block_sectors = roundup(1 << rdev->badblocks.shift,
> + bdev_logical_block_size(rdev->bdev) >> 9);
> sector = r1_bio->sector;
> sectors = ((sector + block_sectors)
> & ~(sector_t)(block_sectors - 1))
> --
> To unsubscribe from this list: send the line "unsubscribe linux-raid" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
Thanks. I've applied this patch and a similar one for RAID10.
This patch had spaces where it should have had tabs (and had no space at all
on one line which should have had a space).
I've fixed all that up, but if you find yourself submitting more patches in
future it would be worth working out how to convince your mailer to send the
patches cleanly with no TAB->space conversions.
Thanks,
NeilBrown
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 811 bytes --]
^ permalink raw reply
* Re: mdadm failed to remove internal bitmap
From: NeilBrown @ 2015-02-16 3:34 UTC (permalink / raw)
To: gary; +Cc: linux-raid
In-Reply-To: <54DDC806.5010803@gmail.com>
[-- Attachment #1: Type: text/plain, Size: 1717 bytes --]
On Fri, 13 Feb 2015 17:46:46 +0800 gary <gary.mdjiang@gmail.com> wrote:
> Hi,
>
> I used v3.3.1 mdadm to do some test for bitmap, but when switch bitmap from
> internal to none, the output shows fail info about remove internal
> bitmap, is it
> just a warning? Since the bitmap seems to be cleared, and it doesn't
> show with
> v3.2.6 mdadm with the same steps.
>
> linux:~ # mdadm --create md0 --raid-devices=2 --level=mirror
> --assume-clean /dev/vdb /dev/vdc
> mdadm: Note: this array has metadata at the start and
> may not be suitable as a boot device. If you plan to
> store '/boot' on this device please ensure that
> your boot-loader understands md/v1.x metadata, or use
> --metadata=0.90
> Continue creating array? y
> mdadm: Defaulting to version 1.2 metadata
> mdadm: array /dev/md/md0 started.
> linux:~ # cat /proc/mdstat
> Personalities : [raid1]
> md127 : active raid1 vdc[1] vdb[0]
> 523712 blocks super 1.2 [2/2] [UU]
>
> unused devices: <none>
> linux:~ # mdadm --grow --bitmap=internal /dev/md127
> linux:~ # cat /proc/mdstat
> Personalities : [raid1]
> md127 : active raid1 vdc[1] vdb[0]
> 523712 blocks super 1.2 [2/2] [UU]
> bitmap: 1/1 pages [4KB], 65536KB chunk
>
> unused devices: <none>
> linux:~ # mdadm --grow --bitmap=none /dev/md127
> mdadm: failed to remove internal bitmap.
> linux:~ # cat /proc/mdstat
> Personalities : [raid1]
> md127 : active raid1 vdc[1] vdb[0]
> 523712 blocks super 1.2 [2/2] [UU]
>
> unused devices: <none>
>
I cannot reproduce this.
What kernel are you running?
Please use "strace" on mdadm in a case where it fails, and post the result.
NeilBrown
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 811 bytes --]
^ permalink raw reply
* Re: re-add POLICY
From: NeilBrown @ 2015-02-16 3:28 UTC (permalink / raw)
To: Chris; +Cc: linux-raid
In-Reply-To: <loom.20150214T222329-324@post.gmane.org>
[-- Attachment #1: Type: text/plain, Size: 1626 bytes --]
On Sat, 14 Feb 2015 21:59:34 +0000 (UTC) Chris <email.bug@arcor.de> wrote:
>
> Hi all,
>
> I'd like mdadm to automatically attempt to re-sync raid members after they
> where temporarily removed from the system.
>
> I would have thought "POLICY domain=default action=re-add" should allow this,
> and found a prior post that also seemed to want/test that behaviour.
> But as I understand the answer given there
> http://permalink.gmane.org/gmane.linux.raid/47516
> mdadm is expected to exit with an error (not re-add) upon plugging the
> device back in?
>
> with:
> mdadm: can only add /dev/loop2 to /dev/md0 as a spare, and force-spare is
> not set.
> mdadm: failed to add /dev/loop2 to existing array /dev/md0: Invalid argument.
>
> For one, I don't understand what the error messages is trying to tell me, about
> an invalid argument that was never supplied to --incremental?
>
> But more importantly, how can priorly diconnected devices (marked failed
> with non-future event count) get re-synced automatically when they are
> plugged in again?
> (avoiding manual mdadm /dev/mdX --add /dev/sdYZ hassle)
>
Does your array have a write-intent bitmap configured?
If it does, then "POLICY action=re-add" really should work.
If it doesn't, then maybe you need "POLICY action=spare".
This isn't the default, because depending on exactly how/why the device
failed, it may not be safe to treat it as a spare.
If the above does not help, please report:
- kernel version
- mdadm version
- "mdadm --examine" output of at least one good drive and one failed drive.
NeilBrown
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 811 bytes --]
^ permalink raw reply
* Re: re-add POLICY: conflict detection?
From: Chris @ 2015-02-15 19:03 UTC (permalink / raw)
To: linux-raid
In-Reply-To: <loom.20150214T222329-324@post.gmane.org>
thinking about the "invalid argument" message...
with "action=re-add":
# mdadm --incremental /dev/loop2
mdadm: can only add /dev/loop2 to /dev/md0 as a spare, and force-spare is
not set.
mdadm: failed to add /dev/loop2 to existing array /dev/md0: Invalid argument.
My guess is that mdadm may not be adding back the failed disk, because it is
unsure wether it may have run separately, and may have newer data on it?
I thought it may be possible to clearly distinguish between clean re-adds
and conflicts, by doing something like this:
* If a member fails (or is missing when starting degraded) write this info
into some failed_at_event_count field belonging to the failed member in the
superblock of every remaining raid member device in the array.
Now, if an array part that got unplugged reappears and still has the event
count that matches the failed_at_event_count that was recorded in the
superblocks of the still running disks, and the reappearing part's
superblock has no failed_at_event_count values for any member of the running
array, the reappearing part is ok to be automatically re-synced.
But if the reappearing disk claims a member of the already running array has
failed, or it reappeared with a different event count than its
faile_at_event_count field in the superblocks of the running array says, a
conflict has arisen and a sync may only be done with manual --force.
Cheers,
Chris
^ permalink raw reply
* Re: update mdadm version to 3.3.2 in Centos 6.4
From: alpha lin @ 2015-02-15 18:48 UTC (permalink / raw)
To: Jes Sorensen; +Cc: linux-raid
In-Reply-To: <wrfjtwz3t9na.fsf@redhat.com>
Hi Jes,
I have my Oracle database running on Centos 6.4. There will be
great if I can stay in Centos 6.4 rather than upgrade to Centos 6.6.
Alpha.
On Tue, Feb 3, 2015 at 9:19 PM, Jes Sorensen <Jes.Sorensen@redhat.com> wrote:
> alpha lin <lin.alpha@gmail.com> writes:
>> Hi All,
>>
>> I would like to know the possibility of update the mdadm package
>> from 3.2.5 to 3.3.2 in Redhat Centos 6.4 x64 system.
>> The problem I have is I found that I can create a raid volume
>> without any problem. But when reboot the system. The md volume become
>> inactive.
>> Below is my step:
>> 1. Check out mdadm package from the github.
>> 2. unzip the mdadm package and execute "make" and "make install"
>> 3. reboot.
>> 4. after system reboot, run mdadm to create a imsm RAID Volume.
>> [root@localhost rules.d]# mdadm --version
>> mdadm - v3.3.2 - 21st August 2014
>
> Any reason why you haven't upgraded to centos 6.6? 6.4 is ancient, and
> 6.6 will at least come with mdadm-3.3.
>
> Jes
>
>
^ permalink raw reply
* re-add POLICY
From: Chris @ 2015-02-14 21:59 UTC (permalink / raw)
To: linux-raid
Hi all,
I'd like mdadm to automatically attempt to re-sync raid members after they
where temporarily removed from the system.
I would have thought "POLICY domain=default action=re-add" should allow this,
and found a prior post that also seemed to want/test that behaviour.
But as I understand the answer given there
http://permalink.gmane.org/gmane.linux.raid/47516
mdadm is expected to exit with an error (not re-add) upon plugging the
device back in?
with:
mdadm: can only add /dev/loop2 to /dev/md0 as a spare, and force-spare is
not set.
mdadm: failed to add /dev/loop2 to existing array /dev/md0: Invalid argument.
For one, I don't understand what the error messages is trying to tell me, about
an invalid argument that was never supplied to --incremental?
But more importantly, how can priorly diconnected devices (marked failed
with non-future event count) get re-synced automatically when they are
plugged in again?
(avoiding manual mdadm /dev/mdX --add /dev/sdYZ hassle)
Cheers,
Chris
^ permalink raw reply
* [PATCH 1/1] Use dev_t for devnm2devid and devid2devnm
From: Mike Lovell @ 2015-02-14 1:08 UTC (permalink / raw)
To: neilb; +Cc: linux-raid, mlovell
In-Reply-To: <1423876088-7169-1-git-send-email-mlovell@bluehost.com>
Commit 4dd2df0966ec added a trip through makedev(), major(), and minor() for
device major and minor numbers. This would cause mdadm to fail in operating
on a device with a minor number bigger than (2^19)-1 due to it changing
from dev_t to a signed int and back.
Where this was found as a problem was when a array was created with a device
specified as a name like /dev/md/raidname and there were already 128 arrays
on the system. In this case, mdadm would chose 1048575 ((2^20)-1) for the
array and minor number. This would cause the major and minor number to become
negative when generated from devnm2devid() and passed to major() and minor()
in open_dev_excl(). open_dev_excl() would then call dev_open() which would
detect the negative minor number and call open() on the *char containing the
major:minor pair which isn't a valid file.
Signed-off-by: Mike Lovell <mlovell@bluehost.com>
---
Detail.c | 4 ++--
Grow.c | 2 +-
lib.c | 2 +-
mapfile.c | 2 +-
mdadm.h | 4 ++--
mdopen.c | 4 ++--
tests/00largemdnumber | 7 +++++++
util.c | 6 +++---
8 files changed, 19 insertions(+), 12 deletions(-)
create mode 100644 tests/00largemdnumber
diff --git a/Detail.c b/Detail.c
index dd72ede..18fa5df 100644
--- a/Detail.c
+++ b/Detail.c
@@ -130,7 +130,7 @@ int Detail(char *dev, struct context *c)
/* This is a subarray of some container.
* We want the name of the container, and the member
*/
- int devid = devnm2devid(st->container_devnm);
+ dev_t devid = devnm2devid(st->container_devnm);
int cfd, err;
member = subarray;
@@ -573,7 +573,7 @@ This is pretty boring
char path[200];
char vbuf[1024];
int nlen = strlen(sra->sys_name);
- int devid;
+ dev_t devid;
if (de->d_name[0] == '.')
continue;
sprintf(path, "/sys/block/%s/md/metadata_version",
diff --git a/Grow.c b/Grow.c
index b78d063..e8f6a2a 100644
--- a/Grow.c
+++ b/Grow.c
@@ -3467,7 +3467,7 @@ int reshape_container(char *container, char *devname,
int fd;
struct mdstat_ent *mdstat;
char *adev;
- int devid;
+ dev_t devid;
sysfs_free(cc);
diff --git a/lib.c b/lib.c
index 6808f62..e9f7018 100644
--- a/lib.c
+++ b/lib.c
@@ -84,7 +84,7 @@ char *devid2kname(int devid)
return NULL;
}
-char *devid2devnm(int devid)
+char *devid2devnm(dev_t devid)
{
char path[30];
char link[200];
diff --git a/mapfile.c b/mapfile.c
index 41599df..9135450 100644
--- a/mapfile.c
+++ b/mapfile.c
@@ -374,7 +374,7 @@ void RebuildMap(void)
char dn[30];
int dfd;
int ok;
- int devid;
+ dev_t devid;
struct supertype *st;
char *subarray = NULL;
char *path;
diff --git a/mdadm.h b/mdadm.h
index 141f963..13279a5 100644
--- a/mdadm.h
+++ b/mdadm.h
@@ -1372,8 +1372,8 @@ extern char *find_free_devnm(int use_partitions);
extern void put_md_name(char *name);
extern char *devid2kname(int devid);
-extern char *devid2devnm(int devid);
-extern int devnm2devid(char *devnm);
+extern char *devid2devnm(dev_t devid);
+extern dev_t devnm2devid(char *devnm);
extern char *get_md_name(char *devnm);
extern char DefaultConfFile[];
diff --git a/mdopen.c b/mdopen.c
index 28410f4..e71d758 100644
--- a/mdopen.c
+++ b/mdopen.c
@@ -348,7 +348,7 @@ int create_mddev(char *dev, char *name, int autof, int trustworthy,
if (lstat(devname, &stb) == 0) {
/* Must be the correct device, else error */
if ((stb.st_mode&S_IFMT) != S_IFBLK ||
- stb.st_rdev != (dev_t)devnm2devid(devnm)) {
+ stb.st_rdev != devnm2devid(devnm)) {
pr_err("%s exists but looks wrong, please fix\n",
devname);
return -1;
@@ -452,7 +452,7 @@ char *find_free_devnm(int use_partitions)
if (!use_udev()) {
/* make sure it is new to /dev too, at least as a
* non-standard */
- int devid = devnm2devid(devnm);
+ dev_t devid = devnm2devid(devnm);
if (devid) {
char *dn = map_dev(major(devid),
minor(devid), 0);
diff --git a/tests/00largemdnumber b/tests/00largemdnumber
new file mode 100644
index 0000000..432e5f4
--- /dev/null
+++ b/tests/00largemdnumber
@@ -0,0 +1,7 @@
+
+# create a simple linear with a large device number
+
+mdadm -CR /dev/md1048575 -l linear -n3 $dev0 $dev1 $dev2
+check linear
+testdev /dev/md1048575 3 $mdsize2_l 1
+mdadm -S /dev/md1048575
diff --git a/util.c b/util.c
index 6f1e2b1..a11f995 100644
--- a/util.c
+++ b/util.c
@@ -779,7 +779,7 @@ int get_data_disks(int level, int layout, int raid_disks)
return data_disks;
}
-int devnm2devid(char *devnm)
+dev_t devnm2devid(char *devnm)
{
/* First look in /sys/block/$DEVNM/dev for %d:%d
* If that fails, try parsing out a number
@@ -916,7 +916,7 @@ int dev_open(char *dev, int flags)
int open_dev_flags(char *devnm, int flags)
{
- int devid;
+ dev_t devid;
char buf[20];
devid = devnm2devid(devnm);
@@ -934,7 +934,7 @@ int open_dev_excl(char *devnm)
char buf[20];
int i;
int flags = O_RDWR;
- int devid = devnm2devid(devnm);
+ dev_t devid = devnm2devid(devnm);
long delay = 1000;
sprintf(buf, "%d:%d", major(devid), minor(devid));
--
1.9.1
^ permalink raw reply related
* [PATCH 0/1] Bug fix for mdadm 3.3 with large md device numbers
From: Mike Lovell @ 2015-02-14 1:08 UTC (permalink / raw)
To: neilb; +Cc: linux-raid, mlovell
A co-worker discovered a situation were new arrays were not being created on
a system. Invoking mdadm similar to 'mdadm --create /dev/md/array-name ...'
would fail where 'mdadm --create /dev/md999 ...' would not. This system had
previously created 128 arrays successfully. He was able to work around the
problem by reverting from mdadm 3.3 to 3.2. The error reported from mdadm
was "mdadm: unexpected failure opening /dev/md1048575."
Some digging with strace showed that mdadm would try to call
open("/sys/block/md1048575/dev") which would fail and then would fail on a
call to open("-4087:-1"). I tested 'mdadm --create /dev/md1048575 ...' on an
empty test VM which would fail as well.
The problem was traced to the md device number being used to create a
major:minor pair which would be passed to makedev(). The result, which is a
dev_t or u32, was then being used as a signed int before being passed to
major() and minor() and then into a string as signed ints. This meant that
the major:minor string had negative numbers in it causing dev_open() to not
recognize it as a valid pair and just calling open() on the string.
The large number was generated because the problem system already had 128
arrays on it. This caused find_free_devnm() to loop from 0 to (1<<20)-1, or
1048575. Triggering the bug is done by specifying a md device number larger
than (1<<19)-1 or by creating a md array by name on a system with 128 already
configured arrays.
Originally, I was going to modify find_free_devnm to loop to (1<<19)-1 but,
since 3.2 works with the larger numbers, I decided to change the signed int
use around devnm2devid and devid2devnm. This patch has been tested against
a number of the tests that weren't already failing on my test system and
didn't cause any more tests to fail. I didn't test all but got the basics.
I am new to the mdadm source and not normally a C developer so this may not
be the best way to fix this but it seems to be working.
^ permalink raw reply
* [PATCH 3/3] md bitmap: export bitmap_destroy() to support dm-raid down takover to raid0
From: heinzm @ 2015-02-13 18:48 UTC (permalink / raw)
To: linux-raid; +Cc: Heinz Mauelshagen
From: Heinz Mauelshagen <heinzm@redhat.com>
This patch exports symbol bitmap_destroy to allow dm-raid to remove
bitmaps when performing a down takeover to md raid0.
Signed-off-by: Heinz Mauelshagen <heinzm@redhat.com>
Tested-by: Heinz Mauelshagen <heinzm@redhat.com>
---
drivers/md/bitmap.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/drivers/md/bitmap.c b/drivers/md/bitmap.c
index 3a57679..b484d15 100644
--- a/drivers/md/bitmap.c
+++ b/drivers/md/bitmap.c
@@ -1631,6 +1631,7 @@ void bitmap_destroy(struct mddev *mddev)
bitmap_free(bitmap);
}
+EXPORT_SYMBOL_GPL(bitmap_destroy);
/*
* initialize the bitmap structure
--
2.1.0
^ permalink raw reply related
* [PATCH 2/3] md raid0: access mddev->queue (request queue member) conditionally because it is not set when accessed from dm-raid
From: heinzm @ 2015-02-13 18:48 UTC (permalink / raw)
To: linux-raid; +Cc: Heinz Mauelshagen
From: Heinz Mauelshagen <heinzm@redhat.com>
The patch makes 3 references to mddev->queue in the raid0 personality
conditional in order to allow for it to be accessed from dm-raid.
Mandatory, because md instances underneath dm-raid don't manage
a request queue of their own which'd lead to oopses without the patch.
Signed-off-by: Heinz Mauelshagen <heinzm@redhat.com>
Tested-by: Heinz Mauelshagen <heinzm@redhat.com>
---
drivers/md/raid0.c | 48 +++++++++++++++++++++++++++---------------------
1 file changed, 27 insertions(+), 21 deletions(-)
diff --git a/drivers/md/raid0.c b/drivers/md/raid0.c
index a13f738..d2e037d 100644
--- a/drivers/md/raid0.c
+++ b/drivers/md/raid0.c
@@ -271,14 +271,16 @@ static int create_strip_zones(struct mddev *mddev, struct r0conf **private_conf)
goto abort;
}
- blk_queue_io_min(mddev->queue, mddev->chunk_sectors << 9);
- blk_queue_io_opt(mddev->queue,
- (mddev->chunk_sectors << 9) * mddev->raid_disks);
-
- if (!discard_supported)
- queue_flag_clear_unlocked(QUEUE_FLAG_DISCARD, mddev->queue);
- else
- queue_flag_set_unlocked(QUEUE_FLAG_DISCARD, mddev->queue);
+ if (mddev->queue) {
+ blk_queue_io_min(mddev->queue, mddev->chunk_sectors << 9);
+ blk_queue_io_opt(mddev->queue,
+ (mddev->chunk_sectors << 9) * mddev->raid_disks);
+
+ if (!discard_supported)
+ queue_flag_clear_unlocked(QUEUE_FLAG_DISCARD, mddev->queue);
+ else
+ queue_flag_set_unlocked(QUEUE_FLAG_DISCARD, mddev->queue);
+ }
pr_debug("md/raid0:%s: done.\n", mdname(mddev));
*private_conf = conf;
@@ -429,9 +431,12 @@ static int raid0_run(struct mddev *mddev)
}
if (md_check_no_bitmap(mddev))
return -EINVAL;
- blk_queue_max_hw_sectors(mddev->queue, mddev->chunk_sectors);
- blk_queue_max_write_same_sectors(mddev->queue, mddev->chunk_sectors);
- blk_queue_max_discard_sectors(mddev->queue, mddev->chunk_sectors);
+
+ if (mddev->queue) {
+ blk_queue_max_hw_sectors(mddev->queue, mddev->chunk_sectors);
+ blk_queue_max_write_same_sectors(mddev->queue, mddev->chunk_sectors);
+ blk_queue_max_discard_sectors(mddev->queue, mddev->chunk_sectors);
+ }
/* if private is not null, we are here after takeover */
if (mddev->private == NULL) {
@@ -448,16 +453,17 @@ static int raid0_run(struct mddev *mddev)
printk(KERN_INFO "md/raid0:%s: md_size is %llu sectors.\n",
mdname(mddev),
(unsigned long long)mddev->array_sectors);
- /* calculate the max read-ahead size.
- * For read-ahead of large files to be effective, we need to
- * readahead at least twice a whole stripe. i.e. number of devices
- * multiplied by chunk size times 2.
- * If an individual device has an ra_pages greater than the
- * chunk size, then we will not drive that device as hard as it
- * wants. We consider this a configuration error: a larger
- * chunksize should be used in that case.
- */
- {
+
+ if (mddev->queue) {
+ /* calculate the max read-ahead size.
+ * For read-ahead of large files to be effective, we need to
+ * readahead at least twice a whole stripe. i.e. number of devices
+ * multiplied by chunk size times 2.
+ * If an individual device has an ra_pages greater than the
+ * chunk size, then we will not drive that device as hard as it
+ * wants. We consider this a configuration error: a larger
+ * chunksize should be used in that case.
+ */
int stripe = mddev->raid_disks *
(mddev->chunk_sectors << 9) / PAGE_SIZE;
if (mddev->queue->backing_dev_info.ra_pages < 2* stripe)
--
2.1.0
^ permalink raw reply related
* [PATCH 1/3] md core: add 2 API functions for takeover and resize to support dm-raid
From: heinzm @ 2015-02-13 18:48 UTC (permalink / raw)
To: linux-raid; +Cc: Heinz Mauelshagen
From: Heinz Mauelshagen <heinzm@redhat.com>
These 2 added external functions allow the device mapper raid target (dm-raid)
to access the md raid takeover and resize funtionality;
reshape API extensions are not needed in lieu of the existing md personality ones.
The patch makes a reference to mddev->queue conditional as well, because
md instances underneath dm-raid don't manage a request queue of their own.
Signed-off-by: Heinz Mauelshagen <heinzm@redhat.com>
Tested-by: Heinz Mauelshagen <heinzm@redhat.com>
---
drivers/md/md.c | 39 ++++++++++++++++++++++++++++++---------
drivers/md/md.h | 3 +++
2 files changed, 33 insertions(+), 9 deletions(-)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index c8d2bac..fb9907c 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -3440,7 +3440,8 @@ level_store(struct mddev *mddev, const char *buf, size_t len)
mddev->in_sync = 1;
del_timer_sync(&mddev->safemode_timer);
}
- blk_set_stacking_limits(&mddev->queue->limits);
+ if (mddev->queue)
+ blk_set_stacking_limits(&mddev->queue->limits);
pers->run(mddev);
set_bit(MD_CHANGE_DEVS, &mddev->flags);
mddev_resume(mddev);
@@ -3454,6 +3455,15 @@ out_unlock:
return rv;
}
+/* API to expose level_store() to dm-raid target */
+int md_takeover(struct mddev *mddev, const char *buf)
+{
+ ssize_t r = level_store(mddev, buf, strlen(buf));
+
+ return r < 0 ? (int) r : 0;
+}
+EXPORT_SYMBOL_GPL(md_takeover);
+
static struct md_sysfs_entry md_level =
__ATTR(level, S_IRUGO|S_IWUSR, level_show, level_store);
@@ -3987,18 +3997,15 @@ size_show(struct mddev *mddev, char *page)
static int update_size(struct mddev *mddev, sector_t num_sectors);
-static ssize_t
-size_store(struct mddev *mddev, const char *buf, size_t len)
+/* API to expose size_store() to dm-raid target */
+int md_resize(struct mddev *mddev, sector_t sectors)
{
+ int err;
+
/* If array is inactive, we can reduce the component size, but
* not increase it (except from 0).
* If array is active, we can try an on-line resize
*/
- sector_t sectors;
- int err = strict_blocks_to_sectors(buf, §ors);
-
- if (err < 0)
- return err;
err = mddev_lock(mddev);
if (err)
return err;
@@ -4013,7 +4020,21 @@ size_store(struct mddev *mddev, const char *buf, size_t len)
err = -ENOSPC;
}
mddev_unlock(mddev);
- return err ? err : len;
+ return err;
+}
+EXPORT_SYMBOL_GPL(md_resize);
+
+/* Compatibility wrapper around md_resize() to keep md internal inbterface */
+static ssize_t
+size_store(struct mddev *mddev, const char *buf, size_t len)
+{
+ sector_t dev_sectors;
+ int err = strict_blocks_to_sectors(buf, &dev_sectors);
+
+ if (!err)
+ err = md_resize(mddev, dev_sectors);
+
+ return err ? (ssize_t) err : len;
}
static struct md_sysfs_entry md_size =
diff --git a/drivers/md/md.h b/drivers/md/md.h
index 318ca8f..892a28a 100644
--- a/drivers/md/md.h
+++ b/drivers/md/md.h
@@ -646,6 +646,9 @@ extern void md_stop_writes(struct mddev *mddev);
extern int md_rdev_init(struct md_rdev *rdev);
extern void md_rdev_clear(struct md_rdev *rdev);
+extern int md_resize(struct mddev *mddev, sector_t dev_sectors);
+extern int md_takeover(struct mddev *mddev, const char *buf);
+
extern void mddev_suspend(struct mddev *mddev);
extern void mddev_resume(struct mddev *mddev);
extern struct bio *bio_clone_mddev(struct bio *bio, gfp_t gfp_mask,
--
2.1.0
^ permalink raw reply related
* [PATCH 0/3] md raid: enhancements to support the device mapper dm-raid target
From: heinzm @ 2015-02-13 18:47 UTC (permalink / raw)
To: linux-raid; +Cc: Heinz Mauelshagen
From: Heinz Mauelshagen <heinzm@redhat.com>
I'm enhancing the device mapper raid target (dm-raid) to take
advantage of so far unused md raid kernel funtionality:
takeover, reshape, resize, addition and removal of devices to/from raid sets.
This series of patches remove constraints doing so.
Patch #1:
add 2 API functions to allow dm-raid to access the raid takeover
and resize functionality (namely md_takeover() and md_resize());
reshape APIs are not needed in lieu of the existing personalilty ones
Patch #2:
because device mapper core manages a request queue per mapped device
utilizing the md make_request API to pass on bios via the dm-raid target,
no md instance underneath it needs to manage a request queue of its own.
Thus dm-raid can't use the md raid0 personality as is, because the latter
accesses the request queue unconditionally in 3 places via mddev->queue
which this patch addresses.
Patch #3:
when dm-raid processes a down takeover to raid0, it needs to destroy
any existing bitmap, because raid0 does not require one. The patch
exports the bitmap_destroy() API to allow dm-raid to remove bitmaps.
Heinz Mauelshagen (3):
md core: add 2 API functions for takeover and resize to support dm-raid
md raid0: access mddev->queue (request queue member) conditionally
because it is not set when accessed from dm-raid
md bitmap: export bitmap_destroy() to support dm-raid down takover to raid0
drivers/md/bitmap.c | 1 +
drivers/md/md.c | 39 ++++++++++++++++++++++++++++++---------
drivers/md/md.h | 3 +++
drivers/md/raid0.c | 48 +++++++++++++++++++++++++++---------------------
4 files changed, 61 insertions(+), 30 deletions(-)
--
2.1.0
^ permalink raw reply
* Re: [PATCH RESEND] Monitor: fix for regression with container devices
From: Artur Paszkiewicz @ 2015-02-13 15:29 UTC (permalink / raw)
To: NeilBrown; +Cc: linux-raid, pawel.baldysiak
In-Reply-To: <20150211153814.333cb17a@notabene.brown>
On 02/11/2015 05:38 AM, NeilBrown wrote:
> On Mon, 9 Feb 2015 11:13:50 +0100 Artur Paszkiewicz
> <artur.paszkiewicz@intel.com> wrote:
>
> > This patch fixes 2 problems introduced by commit 9a518d8: not closing a
> > file descriptor and ignoring container devices. Array state is always
> > "inactive" for containers, so we make sure that the device is not a
> > container by reading also the "level" sysfs entry.
> >
> > Signed-off-by: Artur Paszkiewicz <artur.paszkiewicz@intel.com>
> > Reviewed-by: Pawel Baldysiak <pawel.baldysiak@intel.com>
> > ---
> > Monitor.c | 14 ++++++++++----
> > 1 file changed, 10 insertions(+), 4 deletions(-)
> >
> > diff --git a/Monitor.c b/Monitor.c
> > index 971d2ec..66d67ba 100644
> > --- a/Monitor.c
> > +++ b/Monitor.c
> > @@ -483,11 +483,17 @@ static int check_array(struct state *st, struct mdstat_ent *mdstat,
> > strncmp(buf,"inact",5) == 0) {
> > if (fd >= 0)
> > close(fd);
> > - if (!st->err)
> > - alert("DeviceDisappeared", dev, NULL, ainfo);
> > - st->err++;
> > - return 0;
> > + fd = sysfs_open(st->devnm, NULL, "level");
> > + if (fd < 0 || read(fd, buf, 10) != 0) {
> > + if (fd >= 0)
> > + close(fd);
> > + if (!st->err)
> > + alert("DeviceDisappeared", dev, NULL, ainfo);
> > + st->err++;
> > + return 0;
> > + }
> > }
> > + close(fd);
> > }
> > fd = open(dev, O_RDONLY);
> > if (fd < 0) {
>
> Thanks for the patch.
>
> I don't think I agree with the logic of using 'level' though.
> For the sort of arrays that I need to ignore here, 'level' will be empty.
>
> It would make sense to test 'metadata' though. If that starts 'external:',
> then we don't want to ignore the array.
>
> Could you confirm that this works please?
>
Hi Neil,
I tested your patch. I assume you wanted to use 'metadata_version',
because there is no 'metadata' attribute, right? I had also thought
about that, but simply looking for 'external:' is not enough to
determine that the array is a container - for volumes inside the
container it looks like this: 'external:/md127/0'. But I think that the
arrays you want to ignore will just have 'none' there, so maybe it can
be done like this?
diff --git a/Monitor.c b/Monitor.c
index 971d2ec..83daf3b 100644
--- a/Monitor.c
+++ b/Monitor.c
@@ -483,11 +483,18 @@ static int check_array(struct state *st, struct mdstat_ent *mdstat,
strncmp(buf,"inact",5) == 0) {
if (fd >= 0)
close(fd);
- if (!st->err)
- alert("DeviceDisappeared", dev, NULL, ainfo);
- st->err++;
- return 0;
+ fd = sysfs_open(st->devnm, NULL, "metadata_version");
+ if (fd < 0 || read(fd, buf, 4) < 0 ||
+ strncmp(buf, "none", 4) == 0) {
+ if (fd >= 0)
+ close(fd);
+ if (!st->err)
+ alert("DeviceDisappeared", dev, NULL, ainfo);
+ st->err++;
+ return 0;
+ }
}
+ close(fd);
}
fd = open(dev, O_RDONLY);
if (fd < 0) {
Thanks,
Artur
> Thanks,
> NeilBrown
>
> diff --git a/Monitor.c b/Monitor.c
> index 971d2ecbea72..6e085cb24993 100644
> --- a/Monitor.c
> +++ b/Monitor.c
> @@ -483,11 +483,18 @@ static int check_array(struct state *st, struct mdstat_ent *mdstat,
> strncmp(buf,"inact",5) == 0) {
> if (fd >= 0)
> close(fd);
> - if (!st->err)
> - alert("DeviceDisappeared", dev, NULL, ainfo);
> - st->err++;
> - return 0;
> + fd = sysfs_open(st->devnm, NULL, "metadata");
> + if (fd < 0 || read(fd, buf, 9) != 9 ||
> + strncmp(buf, "external:", 9) != 0) {
> + if (fd >= 0)
> + close(fd);
> + if (!st->err)
> + alert("DeviceDisappeared", dev, NULL, ainfo);
> + st->err++;
> + return 0;
> + }
> }
> + close(fd);
> }
> fd = open(dev, O_RDONLY);
> if (fd < 0) {
>
^ permalink raw reply related
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox