* Why MD often doesn't correct read errors?
From: Ethan Wilson @ 2014-10-02 9:11 UTC (permalink / raw)
To: linux-raid
Hello all,
I am testing a system with a failing disk.
This is an MD raid5 with bitmap, disks are over LSI SAS. A pretty normal
setup.
I can show very long dmesgs in which most read errors are apparently not
corrected. However upper layers such as the filesystem do not complain
either, e.g. the filesystem does not go readonly, and no "read error"
received from userspace. So everything actually works, but I can't
understand why!?
Here is one piece of dmesg in which only at [2204.845894] some errors
get corrected by MD, so I see at least 3 errors before that (at time
729, 1207, 2071) which are apparently ignored by everything:
[ 289.360928] EXT4-fs (dm-0): mounted filesystem with ordered data
mode. Opts: (null)
[ 729.141449] sd 6:0:33:0: [sdah] Unhandled sense code
[ 729.141460] sd 6:0:33:0: [sdah]
[ 729.141463] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[ 729.141466] sd 6:0:33:0: [sdah]
[ 729.141467] Sense Key : Medium Error [current]
[ 729.141471] Info fld=0xcba7e3c
[ 729.141473] sd 6:0:33:0: [sdah]
[ 729.141476] Add. Sense: Unrecovered read error
[ 729.141478] sd 6:0:33:0: [sdah] CDB:
[ 729.141480] Read(10): 28 00 0c ba 7e 00 00 00 a0 00
[ 729.141488] end_request: critical medium error, dev sdah, sector
213548604
[ 781.088413] perf samples too long (2510 > 2500), lowering
kernel.perf_event_max_sample_rate to 50000
[ 1207.475752] sd 6:0:33:0: [sdah] Unhandled sense code
[ 1207.475761] sd 6:0:33:0: [sdah]
[ 1207.475762] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[ 1207.475764] sd 6:0:33:0: [sdah]
[ 1207.475765] Sense Key : Medium Error [current]
[ 1207.475767] Info fld=0xd2d89d2
[ 1207.475769] sd 6:0:33:0: [sdah]
[ 1207.475770] Add. Sense: Unrecovered read error
[ 1207.475772] sd 6:0:33:0: [sdah] CDB:
[ 1207.475773] Read(10): 28 00 0d 2d 88 c0 00 01 98 00
[ 1207.475778] end_request: critical medium error, dev sdah, sector
221088210
[ 2071.445584] sd 6:0:33:0: [sdah] Unhandled sense code
[ 2071.445596] sd 6:0:33:0: [sdah]
[ 2071.445599] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[ 2071.445601] sd 6:0:33:0: [sdah]
[ 2071.445603] Sense Key : Medium Error [current]
[ 2071.445607] Info fld=0xc8fd800
[ 2071.445612] sd 6:0:33:0: [sdah]
[ 2071.445614] Add. Sense: Unrecovered read error
[ 2071.445615] sd 6:0:33:0: [sdah] CDB:
[ 2071.445617] Read(10): 28 00 0c 8f d8 00 00 01 c8 00
[ 2071.445622] end_request: critical medium error, dev sdah, sector
210753536
[ 2201.018508] sd 6:0:33:0: [sdah] Unhandled sense code
[ 2201.018522] sd 6:0:33:0: [sdah]
[ 2201.018525] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[ 2201.018528] sd 6:0:33:0: [sdah]
[ 2201.018530] Sense Key : Medium Error [current]
[ 2201.018534] Info fld=0xc8fb450
[ 2201.018537] sd 6:0:33:0: [sdah]
[ 2201.018546] Add. Sense: Unrecovered read error
[ 2201.018551] sd 6:0:33:0: [sdah] CDB:
[ 2201.018552] Read(10): 28 00 0c 8f b4 48 00 00 38 00
[ 2201.018561] end_request: critical medium error, dev sdah, sector
210744400
[ 2203.651727] sd 6:0:33:0: [sdah] Unhandled sense code
[ 2203.651740] sd 6:0:33:0: [sdah]
[ 2203.651743] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[ 2203.651745] sd 6:0:33:0: [sdah]
[ 2203.651747] Sense Key : Medium Error [current]
[ 2203.651752] Info fld=0xc8fb450
[ 2203.651754] sd 6:0:33:0: [sdah]
[ 2203.651756] Add. Sense: Unrecovered read error
[ 2203.651759] sd 6:0:33:0: [sdah] CDB:
[ 2203.651761] Read(10): 28 00 0c 8f b4 50 00 00 30 00
[ 2203.651769] end_request: critical medium error, dev sdah, sector
210744400
[ 2204.845894] md/raid:md201: read error corrected (8 sectors at 996432
on sdah2)
[ 2204.845912] md/raid:md201: read error corrected (8 sectors at 996440
on sdah2)
[ 2204.845915] md/raid:md201: read error corrected (8 sectors at 996448
on sdah2)
[ 2204.845918] md/raid:md201: read error corrected (8 sectors at 996456
on sdah2)
[ 2204.845920] md/raid:md201: read error corrected (8 sectors at 996464
on sdah2)
[ 2204.845923] md/raid:md201: read error corrected (8 sectors at 996472
on sdah2)
Here is a time in which they get corrected a bit more often, but as you
can see most are still skipped:
[97939.727497] sd 6:0:33:0: [sdah] Unhandled sense code
[97939.727512] sd 6:0:33:0: [sdah]
[97939.727515] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[97939.727518] sd 6:0:33:0: [sdah]
[97939.727520] Sense Key : Medium Error [current]
[97939.727524] Info fld=0xd439400
[97939.727526] sd 6:0:33:0: [sdah]
[97939.727529] Add. Sense: Unrecovered read error
[97939.727531] sd 6:0:33:0: [sdah] CDB:
[97939.727533] Read(10): 28 00 0d 43 94 00 00 00 28 00
[97939.727541] end_request: critical medium error, dev sdah, sector
222532608
[97942.216365] sd 6:0:33:0: [sdah] Unhandled sense code
[97942.216378] sd 6:0:33:0: [sdah]
[97942.216381] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[97942.216382] sd 6:0:33:0: [sdah]
[97942.216384] Sense Key : Medium Error [current]
[97942.216387] Info fld=0xd439400
[97942.216388] sd 6:0:33:0: [sdah]
[97942.216390] Add. Sense: Unrecovered read error
[97942.216391] sd 6:0:33:0: [sdah] CDB:
[97942.216393] Read(10): 28 00 0d 43 94 00 00 00 28 00
[97942.216398] end_request: critical medium error, dev sdah, sector
222532608
[97942.625805] md/raid:md201: read error corrected (8 sectors at
12784640 on sdah2)
[97942.625884] md/raid:md201: read error corrected (8 sectors at
12784648 on sdah2)
[97942.625887] md/raid:md201: read error corrected (8 sectors at
12784656 on sdah2)
[97942.625888] md/raid:md201: read error corrected (8 sectors at
12784664 on sdah2)
[97942.625890] md/raid:md201: read error corrected (8 sectors at
12784672 on sdah2)
[98112.230660] sd 6:0:33:0: [sdah] Unhandled sense code
[98112.230687] sd 6:0:33:0: [sdah]
[98112.230690] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[98112.230692] sd 6:0:33:0: [sdah]
[98112.230694] Sense Key : Medium Error [current]
[98112.230698] Info fld=0xcbaca40
[98112.230700] sd 6:0:33:0: [sdah]
[98112.230703] Add. Sense: Unrecovered read error
[98112.230705] sd 6:0:33:0: [sdah] CDB:
[98112.230707] Read(10): 28 00 0c ba ca 40 00 00 08 00
[98112.230715] end_request: critical medium error, dev sdah, sector
213568064
[99107.714394] sd 6:0:33:0: [sdah] Unhandled sense code
[99107.714443] sd 6:0:33:0: [sdah]
[99107.714444] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[99107.714446] sd 6:0:33:0: [sdah]
[99107.714447] Sense Key : Medium Error [current]
[99107.714450] Info fld=0xcba46c8
[99107.714451] sd 6:0:33:0: [sdah]
[99107.714453] Add. Sense: Unrecovered read error
[99107.714455] sd 6:0:33:0: [sdah] CDB:
[99107.714456] Read(10): 28 00 0c ba 46 c0 00 00 20 00
[99107.714461] end_request: critical medium error, dev sdah, sector
213534408
[99110.123110] sd 6:0:33:0: [sdah] Unhandled sense code
[99110.123167] sd 6:0:33:0: [sdah]
[99110.123170] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[99110.123173] sd 6:0:33:0: [sdah]
[99110.123175] Sense Key : Medium Error [current]
[99110.123179] Info fld=0xcba46c8
[99110.123181] sd 6:0:33:0: [sdah]
[99110.123184] Add. Sense: Unrecovered read error
[99110.123187] sd 6:0:33:0: [sdah] CDB:
[99110.123189] Read(10): 28 00 0c ba 46 c0 00 00 20 00
[99110.123197] end_request: critical medium error, dev sdah, sector
213534408
[99111.169398] md/raid:md201: read error corrected (8 sectors at 3786440
on sdah2)
[99111.169404] md/raid:md201: read error corrected (8 sectors at 3786448
on sdah2)
[99111.169406] md/raid:md201: read error corrected (8 sectors at 3786456
on sdah2)
[101221.285568] mpt2sas0: _scsih_sas_broadcast_primitive_event: enter:
phy number(1), width(16)
[101221.288095] mpt2sas0: _scsih_sas_broadcast_primitive_event: enter:
phy number(1), width(16)
[101221.290937] mpt2sas0: _scsih_sas_broadcast_primitive_event: enter:
phy number(1), width(16)
[101221.293768] mpt2sas0: _scsih_sas_broadcast_primitive_event: enter:
phy number(1), width(16)
[101491.327771] sd 6:0:33:0: [sdah] Unhandled sense code
[101491.327813] sd 6:0:33:0: [sdah]
[101491.327815] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[101491.327817] sd 6:0:33:0: [sdah]
[101491.327819] Sense Key : Medium Error [current]
[101491.327822] Info fld=0xd2d7c1c
[101491.327824] sd 6:0:33:0: [sdah]
[101491.327826] Add. Sense: Unrecovered read error
[101491.327828] sd 6:0:33:0: [sdah] CDB:
[101491.327830] Read(10): 28 00 0d 2d 7c 18 00 00 08 00
[101491.327836] end_request: critical medium error, dev sdah, sector
221084700
[112965.864443] sd 6:0:33:0: [sdah] Unhandled sense code
[112965.864469] sd 6:0:33:0: [sdah]
[112965.864471] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[112965.864474] sd 6:0:33:0: [sdah]
[112965.864476] Sense Key : Medium Error [current]
[112965.864480] Info fld=0xc8e1cb1
[112968.322232] sd 6:0:33:0: [sdah]
[112968.322233] Add. Sense: Unrecovered read error
[112968.322235] sd 6:0:33:0: [sdah] CDB:
[112968.322236] Read(10): 28 00 0c 8e 1c 00 00 00 d8 00
[112968.322241] end_request: critical medium error, dev sdah, sector
210640049
[112969.127941] md/raid:md201: read error corrected (8 sectors at 892080
on sdah2)
[112969.127952] md/raid:md201: read error corrected (8 sectors at 892088
on sdah2)
[112969.127954] md/raid:md201: read error corrected (8 sectors at 892096
on sdah2)
[112969.127955] md/raid:md201: read error corrected (8 sectors at 892104
on sdah2)
[112969.127957] md/raid:md201: read error corrected (8 sectors at 892112
on sdah2)
[113352.100011] sd 6:0:33:0: [sdah] Unhandled sense code
[113352.100068] sd 6:0:33:0: [sdah]
[113352.100071] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[113352.100074] sd 6:0:33:0: [sdah]
[113352.100076] Sense Key : Medium Error [current]
[113352.100080] Info fld=0xc8e8448
[113352.100083] sd 6:0:33:0: [sdah]
[113352.100086] Add. Sense: Unrecovered read error
[113352.100088] sd 6:0:33:0: [sdah] CDB:
[113352.100090] Read(10): 28 00 0c 8e 84 30 00 00 38 00
[113352.100099] end_request: critical medium error, dev sdah, sector
210666568
[113354.850395] sd 6:0:33:0: [sdah] Unhandled sense code
[113354.850404] sd 6:0:33:0: [sdah]
[113354.850406] Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE
[113354.850408] sd 6:0:33:0: [sdah]
[113354.850409] Sense Key : Medium Error [current]
[113354.850412] Info fld=0xc8e8448
[113354.850414] sd 6:0:33:0: [sdah]
[113354.850416] Add. Sense: Unrecovered read error
[113354.850417] sd 6:0:33:0: [sdah] CDB:
[113354.850419] Read(10): 28 00 0c 8e 84 30 00 00 38 00
[113354.850424] end_request: critical medium error, dev sdah, sector
210666568
[113355.387298] md/raid:md201: read error corrected (8 sectors at 918600
on sdah2)
[113355.387303] md/raid:md201: read error corrected (8 sectors at 918608
on sdah2)
[113355.387305] md/raid:md201: read error corrected (8 sectors at 918616
on sdah2)
[113355.387307] md/raid:md201: read error corrected (8 sectors at 918624
on sdah2)
As I wrote above, no error is noticed by userspace, so it actually
works, but I don't know why!?
Thanks for info
EW
^ permalink raw reply
* Re: [dm-devel] dm-raid: add RAID discard support
From: NeilBrown @ 2014-10-02 4:04 UTC (permalink / raw)
To: Mike Snitzer
Cc: Heinz Mauelshagen, device-mapper development, Shaohua Li,
Martin K. Petersen, linux RAID, lkml
In-Reply-To: <20141002120049.58dba551@notabene.brown>
[-- Attachment #1: Type: text/plain, Size: 3570 bytes --]
I plan to submit this to Linus tomorrow, hopefully for 3.7, unless there are
complaints. It is in my for-next branch now.
Thanks,
NeilBrown
From aec6f821ed92fac5ae4f1db50279a3999de5872a Mon Sep 17 00:00:00 2001
From: NeilBrown <neilb@suse.de>
Date: Thu, 2 Oct 2014 13:45:00 +1000
Subject: [PATCH] md/raid5: disable 'DISCARD' by default due to safety
concerns.
It has come to my attention (thanks Martin) that 'discard_zeroes_data'
is only a hint. Some devices in some cases don't do what it
says on the label.
The use of DISCARD in RAID5 depends on reads from discarded regions
being predictably zero. If a write to a previously discarded region
performs a read-modify-write cycle it assumes that the parity block
was consistent with the data blocks. If all were zero, this would
be the case. If some are and some aren't this would not be the case.
This could lead to data corruption after a device failure when
data needs to be reconstructed from the parity.
As we cannot trust 'discard_zeroes_data', ignore it by default
and so disallow DISCARD on all raid4/5/6 arrays.
As many devices are trustworthy, and as there are benefits to using
DISCARD, add a module parameter to over-ride this caution and cause
DISCARD to work if discard_zeroes_data is set.
If a site want to enable DISCARD on some arrays but not on others they
should select DISCARD support at the filesystem level, and set the
raid456 module parameter.
raid456.devices_handle_discard_safely=Y
As this is a data-safety issue, I believe this patch is suitable for
-stable.
DISCARD support for RAID456 was added in 3.7
Cc: Shaohua Li <shli@kernel.org>
Cc: "Martin K. Petersen" <martin.petersen@oracle.com>
Cc: Mike Snitzer <snitzer@redhat.com>
Cc: Heinz Mauelshagen <heinzm@redhat.com>
Cc: stable@vger.kernel.org (3.7+)
Fixes: 620125f2bf8ff0c4969b79653b54d7bcc9d40637
Signed-off-by: NeilBrown <neilb@suse.de>
diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c
index 183588b11fc1..9f0fbecd1eb5 100644
--- a/drivers/md/raid5.c
+++ b/drivers/md/raid5.c
@@ -64,6 +64,10 @@
#define cpu_to_group(cpu) cpu_to_node(cpu)
#define ANY_GROUP NUMA_NO_NODE
+static bool devices_handle_discard_safely = false;
+module_param(devices_handle_discard_safely, bool, 0644);
+MODULE_PARM_DESC(devices_handle_discard_safely,
+ "Set to Y if all devices in each array reliably return zeroes on reads from discarded regions");
static struct workqueue_struct *raid5_wq;
/*
* Stripe cache
@@ -6208,7 +6212,7 @@ static int run(struct mddev *mddev)
mddev->queue->limits.discard_granularity = stripe;
/*
* unaligned part of discard request will be ignored, so can't
- * guarantee discard_zerors_data
+ * guarantee discard_zeroes_data
*/
mddev->queue->limits.discard_zeroes_data = 0;
@@ -6233,6 +6237,18 @@ static int run(struct mddev *mddev)
!bdev_get_queue(rdev->bdev)->
limits.discard_zeroes_data)
discard_supported = false;
+ /* Unfortunately, discard_zeroes_data is not currently
+ * a guarantee - just a hint. So we only allow DISCARD
+ * if the sysadmin has confirmed that only safe devices
+ * are in use by setting a module parameter.
+ */
+ if (!devices_handle_discard_safely) {
+ if (discard_supported) {
+ pr_info("md/raid456: discard support disabled due to uncertainty.\n");
+ pr_info("Set raid456.devices_handle_discard_safely=Y to override.\n");
+ }
+ discard_supported = false;
+ }
}
if (discard_supported &&
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 828 bytes --]
^ permalink raw reply related
* Re: UEFI and mdadm questions.
From: Anthonys Lists @ 2014-10-02 0:17 UTC (permalink / raw)
To: Wilson, Jonathan, Linux-RAID
In-Reply-To: <BLU436-SMTP10330CBEBC0CA1E27853BE498B80@phx.gbl>
On 01/10/2014 17:33, Wilson, Jonathan wrote:
I've just been struggling with exactly this trying to move my gentoo
system on to raid ...
> From what I can tell with UEFI I need to set up a UEFI partition with a
> FAT format.
I believe so, yes.
>
> On my current BIOS system I have a Biosboot 1M, /boot Raid1 200M and /
> Raid 1 40G.
>
> Obviously Grub installs to the mbr, and then installs a bit into
> Biosboot which can read raids, hence it can read and boot from /boot.
Is this grub1, or grub2? I gather grub1 actually CAN'T read raid, which
is why you need to use raid version 0.9. Grub then thinks the disk is
plain ext4 or whatever, and reads it fine. grub2 actually handles raid,
and therefore is happy with newer raids like 1.2
>
> Further, from what I can tell, into the UEFI partition can go either a
> kernel & initramfs with UEFI support, or a "loader" that then loads the
> kernel.
>
> What I am unsure about are...
>
> 1) Can the loader/kernel understand md raid? so the / can be in a bog
> standard raid1 v1.2?
Don't think so. Apparently the kernel can NOT put a raid array together
so you have a catch-22 - the kernel needs the raid in order to start
user-space, but it needs user-space in order to find the raid ... :-(
So you need an initramfs to solve the conundrum. And of course, grub1
doesn't work with uefi :-(
>
> 2) I'm guessing I would no longer need the /boot as that would be
> replaced by what ever was in the UEFI partition?
You do need /boot - not least so grub2 can find its config file (that's
not quite true, but close enough...)
>
> 3) as my "/boot" is currently in a raid 1 my life is simple, should any
> changes occur they are replicated to the drives I have set up as /boot
> raid 1, and I installed the mbr portion of boot loader manually on each
> disk and tested pulling one, and then booting from another.. it
> worked :-)
In the world of grub2/raid1.2 I wish things looked that simple. From my
"badly burned novice" viewpoint, I don't think it is.
> So I would like to keep things, if not simple, at least less likely to
> have problems because I forgot to install duplicates of the UEFI on all
> the disks... so can I use a .90v raid1 on the UEFI partition, then
> format it as fat... so that all copies of what ever is in the UEFI
> partition are replicated across the multiple raid1 disks, instead of
> having to remember to copy what ever is in there manually to ach disk?
Interesting idea ... but I get the impression that once you've set up
the uefi partition and got it working, you shouldn't normally ever need
to go back to it. This is, I think, one of those cases you need to set
up a crib sheet and not do anything without it, as you'll forget from
one time to the next (which is why, I think, you want to use raid :-)
Personally, I *wouldn't* want to use raid, on the assumption that I'm
going to screw it up and don't want a mirror trashing disk2 as I trash
disk1.
Cheers,
Wol
^ permalink raw reply
* [PATCH] dm-log-userspace: fix memory leak on failure path in dm_ulog_tfr_init()
From: Alexey Khoroshilov @ 2014-10-01 20:58 UTC (permalink / raw)
To: Alasdair Kergon, Mike Snitzer
Cc: Alexey Khoroshilov, dm-devel, Neil Brown, linux-raid,
linux-kernel, ldv-project
If cn_add_callback() fails in dm_ulog_tfr_init(), it does not
deallocate prealloced memory but calls cn_del_callback().
It looks like a misprint.
Found by Linux Driver Verification project (linuxtesting.org).
Signed-off-by: Alexey Khoroshilov <khoroshilov@ispras.ru>
---
drivers/md/dm-log-userspace-transfer.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/md/dm-log-userspace-transfer.c b/drivers/md/dm-log-userspace-transfer.c
index b428c0ae63d5..39ad9664d397 100644
--- a/drivers/md/dm-log-userspace-transfer.c
+++ b/drivers/md/dm-log-userspace-transfer.c
@@ -272,7 +272,7 @@ int dm_ulog_tfr_init(void)
r = cn_add_callback(&ulog_cn_id, "dmlogusr", cn_ulog_callback);
if (r) {
- cn_del_callback(&ulog_cn_id);
+ kfree(prealloced_cn_msg);
return r;
}
--
1.9.1
^ permalink raw reply related
* UEFI and mdadm questions.
From: Wilson, Jonathan @ 2014-10-01 16:33 UTC (permalink / raw)
To: Linux-RAID
From what I can tell with UEFI I need to set up a UEFI partition with a
FAT format.
On my current BIOS system I have a Biosboot 1M, /boot Raid1 200M and /
Raid 1 40G.
Obviously Grub installs to the mbr, and then installs a bit into
Biosboot which can read raids, hence it can read and boot from /boot.
Further, from what I can tell, into the UEFI partition can go either a
kernel & initramfs with UEFI support, or a "loader" that then loads the
kernel.
What I am unsure about are...
1) Can the loader/kernel understand md raid? so the / can be in a bog
standard raid1 v1.2?
2) I'm guessing I would no longer need the /boot as that would be
replaced by what ever was in the UEFI partition?
3) as my "/boot" is currently in a raid 1 my life is simple, should any
changes occur they are replicated to the drives I have set up as /boot
raid 1, and I installed the mbr portion of boot loader manually on each
disk and tested pulling one, and then booting from another.. it
worked :-)
So I would like to keep things, if not simple, at least less likely to
have problems because I forgot to install duplicates of the UEFI on all
the disks... so can I use a .90v raid1 on the UEFI partition, then
format it as fat... so that all copies of what ever is in the UEFI
partition are replicated across the multiple raid1 disks, instead of
having to remember to copy what ever is in there manually to ach disk?
Jon.
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Mikael Abrahamsson @ 2014-10-01 7:50 UTC (permalink / raw)
To: Roberto Spadim; +Cc: Linux-RAID
In-Reply-To: <CAH3kUhHkot3doqVZQ0TAot0A3e8i8RFWqsiLL35p1OYsvz5Xkw@mail.gmail.com>
On Tue, 30 Sep 2014, Roberto Spadim wrote:
> any idea/experience and information is wellcome
Please read the treads in the archives regarding this issue. It's been
discussed multiple times.
http://board.issociate.de/thread/508953/Re-Software-RAID-and-TRIM.html is
one of them.
--
Mikael Abrahamsson email: swmike@swm.pp.se
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: David Brown @ 2014-10-01 7:33 UTC (permalink / raw)
To: Roberto Spadim, Linux-RAID
In-Reply-To: <CAH3kUhHkot3doqVZQ0TAot0A3e8i8RFWqsiLL35p1OYsvz5Xkw@mail.gmail.com>
On 30/09/14 16:50, Roberto Spadim wrote:
> hi guys!
> i will use a ssd raid1, i want know if raid1 trim is supported at mdadm
> i will use a 840 evo (or evo pro not selected the right one yet) 500gb
> each, raid1, today database size is 100gb, i think it will grow
> 10gb/year, i had many space...
>
> the point are: madm raid1 trim is supported? or should i use lvm?
> should i partition it with 400gb and leave 100gb untouched? or should
> i use a hdd+ssd and dmcache?
>
>
> :) thanks guys, that's a small enterprise solution, they can't buy
> raid cards and sas harddisk are same price of ssd :)
>
> any idea/experience and information is wellcome
>
With reasonably modern SSD's, don't worry about trim. Make sure you've
got enough over-provisioning (look at the specs for the SSD - and if
necessary, leave a 5-10% of the space on the disk as unpartitioned free
space).
The main point of trim is to return blocks to the list of blocks
available for recycling - if you have plenty of overprovisioning, this
list (along with the list of blocks already erased) is always going to
have plenty in it. Thus trim doesn't add anything useful here.
The second effect of trim is to reduce the number of blocks copied
during garbage collection. As it is unlikely that many extra blocks are
copied, and it is done at off-peak times by the SSD, this makes very
little difference - certainly nothing you would notice in most use-cases.
In other words, you should not notice whether trim is used or not.
Having said that, I believe trim works on raid1 with current kernels.
But make sure you only do off-line trim (fstrim), not the "discard"
mount option in ext4 (or equivalent on other filesystems), as discard
mount will give you noticeably slower performance on many operations.
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Roberto Spadim @ 2014-10-01 2:01 UTC (permalink / raw)
To: Brassow Jonathan; +Cc: Linux-RAID
In-Reply-To: <FA95F377-79D3-4DF7-8B41-D7BB1B706660@redhat.com>
humm interesting , just mdraid1 work with trim :)
other doubt now.. when using ssd+hdd, should trim be enabled at filesystem?
2014-09-30 21:36 GMT-03:00 Brassow Jonathan <jbrassow@redhat.com>:
>
> On Sep 30, 2014, at 9:50 AM, Roberto Spadim wrote:
>> the point are: madm raid1 trim is supported? or should i use lvm?
>> should i partition it with 400gb and leave 100gb untouched? or should
>> i use a hdd+ssd and dmcache?
>
> LVM does not yet support trim on RAID. (Funny you should ask! We've just been talking about enabling this in LVM recently and it should go upstream very soon - following into the various distros a little later.)
>
> brassow
--
Roberto Spadim
SPAEmpresarial
Eng. Automação e Controle
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Brassow Jonathan @ 2014-10-01 0:36 UTC (permalink / raw)
To: Roberto Spadim; +Cc: Linux-RAID
In-Reply-To: <CAH3kUhHkot3doqVZQ0TAot0A3e8i8RFWqsiLL35p1OYsvz5Xkw@mail.gmail.com>
On Sep 30, 2014, at 9:50 AM, Roberto Spadim wrote:
> the point are: madm raid1 trim is supported? or should i use lvm?
> should i partition it with 400gb and leave 100gb untouched? or should
> i use a hdd+ssd and dmcache?
LVM does not yet support trim on RAID. (Funny you should ask! We've just been talking about enabling this in LVM recently and it should go upstream very soon - following into the various distros a little later.)
brassow
^ permalink raw reply
* Re: Raid5 hang in 3.14.19
From: NeilBrown @ 2014-09-30 22:54 UTC (permalink / raw)
To: BillStuff; +Cc: linux-raid
In-Reply-To: <542B1EDC.7060803@sbcglobal.net>
[-- Attachment #1: Type: text/plain, Size: 13903 bytes --]
On Tue, 30 Sep 2014 16:21:32 -0500 BillStuff <billstuff2001@sbcglobal.net>
wrote:
> On 09/29/2014 04:59 PM, NeilBrown wrote:
> > On Sun, 28 Sep 2014 23:28:17 -0500 BillStuff <billstuff2001@sbcglobal.net>
> > wrote:
> >
> >> On 09/28/2014 11:08 PM, NeilBrown wrote:
> >>> On Sun, 28 Sep 2014 22:56:19 -0500 BillStuff <billstuff2001@sbcglobal.net>
> >>> wrote:
> >>>
> >>>> On 09/28/2014 09:25 PM, NeilBrown wrote:
> >>>>> On Fri, 26 Sep 2014 17:33:58 -0500 BillStuff <billstuff2001@sbcglobal.net>
> >>>>> wrote:
> >>>>>
> >>>>>> Hi Neil,
> >>>>>>
> >>>>>> I found something that looks similar to the problem described in
> >>>>>> "Re: seems like a deadlock in workqueue when md do a flush" from Sept 14th.
> >>>>>>
> >>>>>> It's on 3.14.19 with 7 recent patches for fixing raid1 recovery hangs.
> >>>>>>
> >>>>>> on this array:
> >>>>>> md3 : active raid5 sdf1[5] sde1[4] sdd1[3] sdc1[2] sdb1[1] sda1[0]
> >>>>>> 104171200 blocks level 5, 64k chunk, algorithm 2 [6/6] [UUUUUU]
> >>>>>> bitmap: 1/5 pages [4KB], 2048KB chunk
> >>>>>>
> >>>>>> I was running a test doing parallel kernel builds, read/write loops, and
> >>>>>> disk add / remove / check loops,
> >>>>>> on both this array and a raid1 array.
> >>>>>>
> >>>>>> I was trying to stress test your recent raid1 fixes, which went well,
> >>>>>> but then after 5 days,
> >>>>>> the raid5 array hung up with this in dmesg:
> >>>>> I think this is different to the workqueue problem you mentioned, though as I
> >>>>> don't know exactly what caused either I cannot be certain.
> >>>>>
> >>>>> From the data you provided it looks like everything is waiting on
> >>>>> get_active_stripe(), or on a process that is waiting on that.
> >>>>> That seems pretty common whenever anything goes wrong in raid5 :-(
> >>>>>
> >>>>> The md3_raid5 task is listed as blocked, but not stack trace is given.
> >>>>> If the machine is still in the state, then
> >>>>>
> >>>>> cat /proc/1698/stack
> >>>>>
> >>>>> might be useful.
> >>>>> (echo t > /proc/sysrq-trigger is always a good idea)
> >>>> Might this help? I believe the array was doing a "check" when things
> >>>> hung up.
> >>> It looks like it was trying to start doing a 'check'.
> >>> The 'resync' thread hadn't been started yet.
> >>> What is 'kthreadd' doing?
> >>> My guess is that it is in try_to_free_pages() waiting for writeout
> >>> for some xfs file page onto the md array ... which won't progress until
> >>> the thread gets started.
> >>>
> >>> That would suggest that we need an async way to start threads...
> >>>
> >>> Thanks,
> >>> NeilBrown
> >>>
> >> I suspect your guess is correct:
> > Thanks for the confirmation.
> >
> > I'm thinking of something like that. Very basic suggestion suggests it
> > instantly crash.
> >
> > If you were to apply this patch and run your test for a week or two, that
> > would increase my confidence (though of course testing doesn't prove the
> > absence of bugs....)
> >
> > Thanks,
> > NeilBrown
> >
>
> Neil,
>
> It locked up already, but in a different way.
> In this case it was triggered by a re-add in a fail / remove / add loop.
>
> It kinda seems like md3_raid5 tried to "collect" (reap) md3_resync
> before it got fully started.
>
>
> INFO: task md3_raid5:1698 blocked for more than 120 seconds.
> Tainted: P O 3.14.19fe-dirty #3
> "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
> e8371dc4 00000046 c106d92d ea483770 d18a9994 ffd2177f ffffffff e8371dac
> c17d6700 c17d6700 e9d3bcc0 c231bcc0 00000420 e8371d98 00007400 00000088
> 00000000 00000000 c33ead90 000003be 00000000 00000005 00000000 0000528b
> Call Trace:
> [<c106d92d>] ? __enqueue_entity+0x6d/0x80
> [<c1072683>] ? enqueue_task_fair+0x2d3/0x660
> [<c153e893>] schedule+0x23/0x60
> [<c153dc25>] schedule_timeout+0x145/0x1c0
> [<c1069cb0>] ? default_wake_function+0x10/0x20
> [<c1065698>] ? update_rq_clock.part.92+0x18/0x50
> [<c1067a65>] ? check_preempt_curr+0x65/0x90
> [<c1067aa8>] ? ttwu_do_wakeup+0x18/0x120
> [<c153effb>] wait_for_common+0x9b/0x110
> [<c1069ca0>] ? wake_up_process+0x40/0x40
> [<c153f087>] wait_for_completion+0x17/0x20
> [<c105b241>] kthread_stop+0x41/0xb0
> [<c14540e1>] md_unregister_thread+0x31/0x40
> [<c145a799>] md_reap_sync_thread+0x19/0x140
> [<c145aba8>] md_check_recovery+0xe8/0x480
> [<f3dc3a10>] raid5d+0x20/0x4c0 [raid456]
> [<c104a022>] ? try_to_del_timer_sync+0x42/0x60
> [<c104a081>] ? del_timer_sync+0x41/0x50
> [<c153dbdd>] ? schedule_timeout+0xfd/0x1c0
> [<c10497c0>] ? detach_if_pending+0xa0/0xa0
> [<c10797b4>] ? finish_wait+0x44/0x60
> [<c1454098>] md_thread+0xe8/0x100
> [<c1079990>] ? __wake_up_sync+0x20/0x20
> [<c1453fb0>] ? md_start_sync+0xb0/0xb0
> [<c105ae21>] kthread+0xa1/0xc0
> [<c15418b7>] ret_from_kernel_thread+0x1b/0x28
> [<c105ad80>] ? kthread_create_on_node+0x110/0x110
> INFO: task kworker/0:0:12491 blocked for more than 120 seconds.
> Tainted: P O 3.14.19fe-dirty #3
> "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
> e6f2ba58 00000046 00000086 00000086 e6f2ba08 c10795ff f395c035 000030ea
> c17d6700 c17d6700 d18a9440 e8973cc0 e6f2ba14 c144fa3d 00000000 e6f2ba2c
> c2054600 f3dbc81b ea5f4c80 c9f64060 c2054600 00000001 c9f64060 e6f2ba5c
> Call Trace:
> [<c10795ff>] ? __wake_up+0x3f/0x50
> [<c144fa3d>] ? md_wakeup_thread+0x2d/0x30
> [<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
> [<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
> [<c153e893>] schedule+0x23/0x60
> [<c153dc25>] schedule_timeout+0x145/0x1c0
> [<c12b1192>] ? blk_finish_plug+0x12/0x40
> [<f3c919f7>] ? _xfs_buf_ioapply+0x287/0x300 [xfs]
> [<c153effb>] wait_for_common+0x9b/0x110
> [<c1069ca0>] ? wake_up_process+0x40/0x40
> [<c153f087>] wait_for_completion+0x17/0x20
> [<f3c91e80>] xfs_buf_iowait+0x50/0xb0 [xfs]
> [<f3c91f17>] ? _xfs_buf_read+0x37/0x40 [xfs]
> [<f3c91f17>] _xfs_buf_read+0x37/0x40 [xfs]
> [<f3c91fa5>] xfs_buf_read_map+0x85/0xe0 [xfs]
> [<f3cefa99>] xfs_trans_read_buf_map+0x1e9/0x3e0 [xfs]
> [<f3caa977>] xfs_alloc_read_agfl+0x97/0xc0 [xfs]
> [<f3cad8a3>] xfs_alloc_fix_freelist+0x193/0x430 [xfs]
> [<f3ce83fe>] ? xlog_grant_sub_space.isra.4+0x1e/0x70 [xfs]
> [<f3cae17e>] xfs_free_extent+0x8e/0x100 [xfs]
> [<c111be1e>] ? kmem_cache_alloc+0xae/0xf0
> [<f3c8e12c>] xfs_bmap_finish+0x13c/0x190 [xfs]
> [<f3cda6f2>] xfs_itruncate_extents+0x1b2/0x2b0 [xfs]
> [<f3c8f0c3>] xfs_free_eofblocks+0x243/0x2e0 [xfs]
> [<f3c9abc6>] xfs_inode_free_eofblocks+0x96/0x140 [xfs]
> [<f3c9ab30>] ? xfs_inode_clear_eofblocks_tag+0x160/0x160 [xfs]
> [<f3c9956d>] xfs_inode_ag_walk.isra.7+0x1cd/0x2f0 [xfs]
> [<f3c9ab30>] ? xfs_inode_clear_eofblocks_tag+0x160/0x160 [xfs]
> [<c12d1631>] ? radix_tree_gang_lookup_tag+0x71/0xb0
> [<f3ce55fb>] ? xfs_perag_get_tag+0x2b/0xb0 [xfs]
> [<f3c9a56d>] ? xfs_inode_ag_iterator_tag+0x3d/0xa0 [xfs]
> [<f3c9a59a>] xfs_inode_ag_iterator_tag+0x6a/0xa0 [xfs]
> [<f3c9ab30>] ? xfs_inode_clear_eofblocks_tag+0x160/0x160 [xfs]
> [<f3c9a82e>] xfs_icache_free_eofblocks+0x2e/0x40 [xfs]
> [<f3c9a858>] xfs_eofblocks_worker+0x18/0x30 [xfs]
> [<c105521c>] process_one_work+0x10c/0x340
> [<c1055d31>] worker_thread+0x101/0x330
> [<c1055c30>] ? manage_workers.isra.27+0x250/0x250
> [<c105ae21>] kthread+0xa1/0xc0
> [<c15418b7>] ret_from_kernel_thread+0x1b/0x28
> [<c105ad80>] ? kthread_create_on_node+0x110/0x110
> INFO: task xfs_fsr:24262 blocked for more than 120 seconds.
> Tainted: P O 3.14.19fe-dirty #3
> "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
> e4d17aec 00000082 00000096 00000096 e4d17a9c c10795ff 00000000 00000000
> c17d6700 c17d6700 e9d3b7b0 c70a8000 e4d17aa8 c144fa3d 00000000 e4d17ac0
> c2054600 f3dbc81b ea7ace80 d7b74cc0 c2054600 00000002 d7b74cc0 e4d17af0
> Call Trace:
> [<c10795ff>] ? __wake_up+0x3f/0x50
> [<c144fa3d>] ? md_wakeup_thread+0x2d/0x30
> [<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
> [<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
> [<c153e893>] schedule+0x23/0x60
> [<c153dc25>] schedule_timeout+0x145/0x1c0
> [<c12b1192>] ? blk_finish_plug+0x12/0x40
> [<f3c919f7>] ? _xfs_buf_ioapply+0x287/0x300 [xfs]
> [<c153effb>] wait_for_common+0x9b/0x110
> [<c1069ca0>] ? wake_up_process+0x40/0x40
> [<c153f087>] wait_for_completion+0x17/0x20
> [<f3c91e80>] xfs_buf_iowait+0x50/0xb0 [xfs]
> [<f3c91f17>] ? _xfs_buf_read+0x37/0x40 [xfs]
> [<f3c91f17>] _xfs_buf_read+0x37/0x40 [xfs]
> [<f3c91fa5>] xfs_buf_read_map+0x85/0xe0 [xfs]
> [<f3cefa29>] xfs_trans_read_buf_map+0x179/0x3e0 [xfs]
> [<f3cde727>] xfs_imap_to_bp+0x67/0xe0 [xfs]
> [<f3cdec7d>] xfs_iread+0x7d/0x3e0 [xfs]
> [<f3c996e8>] ? xfs_inode_alloc+0x58/0x1b0 [xfs]
> [<f3c9a092>] xfs_iget+0x192/0x580 [xfs]
> [<f3caa219>] ? kmem_free+0x19/0x50 [xfs]
> [<f3ca0ce6>] xfs_bulkstat_one_int+0x86/0x2c0 [xfs]
> [<f3ca0f54>] xfs_bulkstat_one+0x34/0x40 [xfs]
> [<f3ca0c10>] ? xfs_internal_inum+0xa0/0xa0 [xfs]
> [<f3ca1360>] xfs_bulkstat+0x400/0x870 [xfs]
> [<f3c9ae86>] xfs_ioc_bulkstat+0xb6/0x160 [xfs]
> [<f3ca0f20>] ? xfs_bulkstat_one_int+0x2c0/0x2c0 [xfs]
> [<f3c9ccb0>] ? xfs_ioc_swapext+0x160/0x160 [xfs]
> [<f3c9d433>] xfs_file_ioctl+0x783/0xa60 [xfs]
> [<c106bf9b>] ? __update_cpu_load+0xab/0xd0
> [<c105dc58>] ? hrtimer_forward+0xa8/0x1b0
> [<c12d4050>] ? timerqueue_add+0x50/0xb0
> [<c108d143>] ? ktime_get+0x53/0xe0
> [<c1094035>] ? clockevents_program_event+0x95/0x130
> [<f3c9ccb0>] ? xfs_ioc_swapext+0x160/0x160 [xfs]
> [<c11349b2>] do_vfs_ioctl+0x2e2/0x4c0
> [<c1095699>] ? tick_program_event+0x29/0x30
> [<c105e42c>] ? hrtimer_interrupt+0x13c/0x2a0
> [<c12d8dfb>] ? lockref_put_or_lock+0xb/0x30
> [<c1134c08>] SyS_ioctl+0x78/0x80
> [<c1540fa8>] syscall_call+0x7/0x7
> INFO: task md3_resync:24817 blocked for more than 120 seconds.
> Tainted: P O 3.14.19fe-dirty #3
> "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
> c2d95e08 00000046 00000092 00000000 00000020 00000092 1befb1af 000030e0
> c17d6700 c17d6700 e9d3c6e0 d18a9950 c2d95ddc c12adbc5 00000000 00000001
> 00000000 00000001 c2d95dec e95f0000 00000001 c2d95e08 00000292 00000292
> Call Trace:
> [<c12adbc5>] ? queue_unplugged+0x45/0x90
> [<c10798bb>] ? prepare_to_wait_event+0x6b/0xd0
> [<c153e893>] schedule+0x23/0x60
> [<c145782d>] md_do_sync+0x8dd/0x1010
> [<c1079990>] ? __wake_up_sync+0x20/0x20
> [<c1453fb0>] ? md_start_sync+0xb0/0xb0
> [<c1454098>] md_thread+0xe8/0x100
> [<c107945f>] ? __wake_up_locked+0x1f/0x30
> [<c1453fb0>] ? md_start_sync+0xb0/0xb0
> [<c105ae21>] kthread+0xa1/0xc0
> [<c15418b7>] ret_from_kernel_thread+0x1b/0x28
> [<c105ad80>] ? kthread_create_on_node+0x110/0x110
>
> Again, the blocked task info is edited for space; hopefully these are
> the important ones,
> but I've got more if it helps.
Thanks for the testing! You have included enough information.
I didn't really like that 'sync_starting' variable when I wrote the patch,
but it seemed do the right thing. It doesn't.
If md_check_recovery() runs again immediately after scheduling the sync
thread to run, it will not have set sync_starting but will find ->sync_thread
is NULL and so will clear MD_RECOVERY_RUNNING. The next time it runs, that
flag is still clear and ->sync_thread is not NULL so it will try to stop the
thread, which deadlocks.
This patch on top of what you have should fix it... but I might end up
redoing the logic a bit.
Thanks,
NeilBrown
diff --git a/drivers/md/md.c b/drivers/md/md.c
index 211a615ae415..3df888b80e76 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -7850,7 +7850,6 @@ void md_check_recovery(struct mddev *mddev)
if (mddev_trylock(mddev)) {
int spares = 0;
- bool sync_starting = false;
if (mddev->ro) {
/* On a read-only array we can:
@@ -7914,7 +7913,7 @@ void md_check_recovery(struct mddev *mddev)
if (!test_and_clear_bit(MD_RECOVERY_NEEDED, &mddev->recovery) ||
test_bit(MD_RECOVERY_FROZEN, &mddev->recovery))
- goto unlock;
+ goto not_running;
/* no recovery is running.
* remove any failed drives, then
* add spares if possible.
@@ -7926,7 +7925,7 @@ void md_check_recovery(struct mddev *mddev)
if (mddev->pers->check_reshape == NULL ||
mddev->pers->check_reshape(mddev) != 0)
/* Cannot proceed */
- goto unlock;
+ goto not_running;
set_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
clear_bit(MD_RECOVERY_RECOVER, &mddev->recovery);
} else if ((spares = remove_and_add_spares(mddev, NULL))) {
@@ -7939,7 +7938,7 @@ void md_check_recovery(struct mddev *mddev)
clear_bit(MD_RECOVERY_RECOVER, &mddev->recovery);
} else if (!test_bit(MD_RECOVERY_SYNC, &mddev->recovery))
/* nothing to be done ... */
- goto unlock;
+ goto not_running;
if (mddev->pers->sync_request) {
if (spares) {
@@ -7951,18 +7950,18 @@ void md_check_recovery(struct mddev *mddev)
}
INIT_WORK(&mddev->del_work, md_start_sync);
queue_work(md_misc_wq, &mddev->del_work);
- sync_starting = true;
+ goto unlock;
}
- unlock:
- wake_up(&mddev->sb_wait);
-
- if (!mddev->sync_thread && !sync_starting) {
+ not_running:
+ if (!mddev->sync_thread) {
clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
if (test_and_clear_bit(MD_RECOVERY_RECOVER,
&mddev->recovery))
if (mddev->sysfs_action)
sysfs_notify_dirent_safe(mddev->sysfs_action);
}
+ unlock:
+ wake_up(&mddev->sb_wait);
mddev_unlock(mddev);
}
}
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 828 bytes --]
^ permalink raw reply related
* Re: Raid5 hang in 3.14.19
From: BillStuff @ 2014-09-30 21:21 UTC (permalink / raw)
To: NeilBrown; +Cc: linux-raid
In-Reply-To: <20140930075950.1d1e3865@notabene.brown>
On 09/29/2014 04:59 PM, NeilBrown wrote:
> On Sun, 28 Sep 2014 23:28:17 -0500 BillStuff <billstuff2001@sbcglobal.net>
> wrote:
>
>> On 09/28/2014 11:08 PM, NeilBrown wrote:
>>> On Sun, 28 Sep 2014 22:56:19 -0500 BillStuff <billstuff2001@sbcglobal.net>
>>> wrote:
>>>
>>>> On 09/28/2014 09:25 PM, NeilBrown wrote:
>>>>> On Fri, 26 Sep 2014 17:33:58 -0500 BillStuff <billstuff2001@sbcglobal.net>
>>>>> wrote:
>>>>>
>>>>>> Hi Neil,
>>>>>>
>>>>>> I found something that looks similar to the problem described in
>>>>>> "Re: seems like a deadlock in workqueue when md do a flush" from Sept 14th.
>>>>>>
>>>>>> It's on 3.14.19 with 7 recent patches for fixing raid1 recovery hangs.
>>>>>>
>>>>>> on this array:
>>>>>> md3 : active raid5 sdf1[5] sde1[4] sdd1[3] sdc1[2] sdb1[1] sda1[0]
>>>>>> 104171200 blocks level 5, 64k chunk, algorithm 2 [6/6] [UUUUUU]
>>>>>> bitmap: 1/5 pages [4KB], 2048KB chunk
>>>>>>
>>>>>> I was running a test doing parallel kernel builds, read/write loops, and
>>>>>> disk add / remove / check loops,
>>>>>> on both this array and a raid1 array.
>>>>>>
>>>>>> I was trying to stress test your recent raid1 fixes, which went well,
>>>>>> but then after 5 days,
>>>>>> the raid5 array hung up with this in dmesg:
>>>>> I think this is different to the workqueue problem you mentioned, though as I
>>>>> don't know exactly what caused either I cannot be certain.
>>>>>
>>>>> From the data you provided it looks like everything is waiting on
>>>>> get_active_stripe(), or on a process that is waiting on that.
>>>>> That seems pretty common whenever anything goes wrong in raid5 :-(
>>>>>
>>>>> The md3_raid5 task is listed as blocked, but not stack trace is given.
>>>>> If the machine is still in the state, then
>>>>>
>>>>> cat /proc/1698/stack
>>>>>
>>>>> might be useful.
>>>>> (echo t > /proc/sysrq-trigger is always a good idea)
>>>> Might this help? I believe the array was doing a "check" when things
>>>> hung up.
>>> It looks like it was trying to start doing a 'check'.
>>> The 'resync' thread hadn't been started yet.
>>> What is 'kthreadd' doing?
>>> My guess is that it is in try_to_free_pages() waiting for writeout
>>> for some xfs file page onto the md array ... which won't progress until
>>> the thread gets started.
>>>
>>> That would suggest that we need an async way to start threads...
>>>
>>> Thanks,
>>> NeilBrown
>>>
>> I suspect your guess is correct:
> Thanks for the confirmation.
>
> I'm thinking of something like that. Very basic suggestion suggests it
> instantly crash.
>
> If you were to apply this patch and run your test for a week or two, that
> would increase my confidence (though of course testing doesn't prove the
> absence of bugs....)
>
> Thanks,
> NeilBrown
>
Neil,
It locked up already, but in a different way.
In this case it was triggered by a re-add in a fail / remove / add loop.
It kinda seems like md3_raid5 tried to "collect" (reap) md3_resync
before it got fully started.
INFO: task md3_raid5:1698 blocked for more than 120 seconds.
Tainted: P O 3.14.19fe-dirty #3
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
e8371dc4 00000046 c106d92d ea483770 d18a9994 ffd2177f ffffffff e8371dac
c17d6700 c17d6700 e9d3bcc0 c231bcc0 00000420 e8371d98 00007400 00000088
00000000 00000000 c33ead90 000003be 00000000 00000005 00000000 0000528b
Call Trace:
[<c106d92d>] ? __enqueue_entity+0x6d/0x80
[<c1072683>] ? enqueue_task_fair+0x2d3/0x660
[<c153e893>] schedule+0x23/0x60
[<c153dc25>] schedule_timeout+0x145/0x1c0
[<c1069cb0>] ? default_wake_function+0x10/0x20
[<c1065698>] ? update_rq_clock.part.92+0x18/0x50
[<c1067a65>] ? check_preempt_curr+0x65/0x90
[<c1067aa8>] ? ttwu_do_wakeup+0x18/0x120
[<c153effb>] wait_for_common+0x9b/0x110
[<c1069ca0>] ? wake_up_process+0x40/0x40
[<c153f087>] wait_for_completion+0x17/0x20
[<c105b241>] kthread_stop+0x41/0xb0
[<c14540e1>] md_unregister_thread+0x31/0x40
[<c145a799>] md_reap_sync_thread+0x19/0x140
[<c145aba8>] md_check_recovery+0xe8/0x480
[<f3dc3a10>] raid5d+0x20/0x4c0 [raid456]
[<c104a022>] ? try_to_del_timer_sync+0x42/0x60
[<c104a081>] ? del_timer_sync+0x41/0x50
[<c153dbdd>] ? schedule_timeout+0xfd/0x1c0
[<c10497c0>] ? detach_if_pending+0xa0/0xa0
[<c10797b4>] ? finish_wait+0x44/0x60
[<c1454098>] md_thread+0xe8/0x100
[<c1079990>] ? __wake_up_sync+0x20/0x20
[<c1453fb0>] ? md_start_sync+0xb0/0xb0
[<c105ae21>] kthread+0xa1/0xc0
[<c15418b7>] ret_from_kernel_thread+0x1b/0x28
[<c105ad80>] ? kthread_create_on_node+0x110/0x110
INFO: task kworker/0:0:12491 blocked for more than 120 seconds.
Tainted: P O 3.14.19fe-dirty #3
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
e6f2ba58 00000046 00000086 00000086 e6f2ba08 c10795ff f395c035 000030ea
c17d6700 c17d6700 d18a9440 e8973cc0 e6f2ba14 c144fa3d 00000000 e6f2ba2c
c2054600 f3dbc81b ea5f4c80 c9f64060 c2054600 00000001 c9f64060 e6f2ba5c
Call Trace:
[<c10795ff>] ? __wake_up+0x3f/0x50
[<c144fa3d>] ? md_wakeup_thread+0x2d/0x30
[<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
[<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
[<c153e893>] schedule+0x23/0x60
[<c153dc25>] schedule_timeout+0x145/0x1c0
[<c12b1192>] ? blk_finish_plug+0x12/0x40
[<f3c919f7>] ? _xfs_buf_ioapply+0x287/0x300 [xfs]
[<c153effb>] wait_for_common+0x9b/0x110
[<c1069ca0>] ? wake_up_process+0x40/0x40
[<c153f087>] wait_for_completion+0x17/0x20
[<f3c91e80>] xfs_buf_iowait+0x50/0xb0 [xfs]
[<f3c91f17>] ? _xfs_buf_read+0x37/0x40 [xfs]
[<f3c91f17>] _xfs_buf_read+0x37/0x40 [xfs]
[<f3c91fa5>] xfs_buf_read_map+0x85/0xe0 [xfs]
[<f3cefa99>] xfs_trans_read_buf_map+0x1e9/0x3e0 [xfs]
[<f3caa977>] xfs_alloc_read_agfl+0x97/0xc0 [xfs]
[<f3cad8a3>] xfs_alloc_fix_freelist+0x193/0x430 [xfs]
[<f3ce83fe>] ? xlog_grant_sub_space.isra.4+0x1e/0x70 [xfs]
[<f3cae17e>] xfs_free_extent+0x8e/0x100 [xfs]
[<c111be1e>] ? kmem_cache_alloc+0xae/0xf0
[<f3c8e12c>] xfs_bmap_finish+0x13c/0x190 [xfs]
[<f3cda6f2>] xfs_itruncate_extents+0x1b2/0x2b0 [xfs]
[<f3c8f0c3>] xfs_free_eofblocks+0x243/0x2e0 [xfs]
[<f3c9abc6>] xfs_inode_free_eofblocks+0x96/0x140 [xfs]
[<f3c9ab30>] ? xfs_inode_clear_eofblocks_tag+0x160/0x160 [xfs]
[<f3c9956d>] xfs_inode_ag_walk.isra.7+0x1cd/0x2f0 [xfs]
[<f3c9ab30>] ? xfs_inode_clear_eofblocks_tag+0x160/0x160 [xfs]
[<c12d1631>] ? radix_tree_gang_lookup_tag+0x71/0xb0
[<f3ce55fb>] ? xfs_perag_get_tag+0x2b/0xb0 [xfs]
[<f3c9a56d>] ? xfs_inode_ag_iterator_tag+0x3d/0xa0 [xfs]
[<f3c9a59a>] xfs_inode_ag_iterator_tag+0x6a/0xa0 [xfs]
[<f3c9ab30>] ? xfs_inode_clear_eofblocks_tag+0x160/0x160 [xfs]
[<f3c9a82e>] xfs_icache_free_eofblocks+0x2e/0x40 [xfs]
[<f3c9a858>] xfs_eofblocks_worker+0x18/0x30 [xfs]
[<c105521c>] process_one_work+0x10c/0x340
[<c1055d31>] worker_thread+0x101/0x330
[<c1055c30>] ? manage_workers.isra.27+0x250/0x250
[<c105ae21>] kthread+0xa1/0xc0
[<c15418b7>] ret_from_kernel_thread+0x1b/0x28
[<c105ad80>] ? kthread_create_on_node+0x110/0x110
INFO: task xfs_fsr:24262 blocked for more than 120 seconds.
Tainted: P O 3.14.19fe-dirty #3
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
e4d17aec 00000082 00000096 00000096 e4d17a9c c10795ff 00000000 00000000
c17d6700 c17d6700 e9d3b7b0 c70a8000 e4d17aa8 c144fa3d 00000000 e4d17ac0
c2054600 f3dbc81b ea7ace80 d7b74cc0 c2054600 00000002 d7b74cc0 e4d17af0
Call Trace:
[<c10795ff>] ? __wake_up+0x3f/0x50
[<c144fa3d>] ? md_wakeup_thread+0x2d/0x30
[<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
[<f3dbc81b>] ? raid5_unplug+0xbb/0x120 [raid456]
[<c153e893>] schedule+0x23/0x60
[<c153dc25>] schedule_timeout+0x145/0x1c0
[<c12b1192>] ? blk_finish_plug+0x12/0x40
[<f3c919f7>] ? _xfs_buf_ioapply+0x287/0x300 [xfs]
[<c153effb>] wait_for_common+0x9b/0x110
[<c1069ca0>] ? wake_up_process+0x40/0x40
[<c153f087>] wait_for_completion+0x17/0x20
[<f3c91e80>] xfs_buf_iowait+0x50/0xb0 [xfs]
[<f3c91f17>] ? _xfs_buf_read+0x37/0x40 [xfs]
[<f3c91f17>] _xfs_buf_read+0x37/0x40 [xfs]
[<f3c91fa5>] xfs_buf_read_map+0x85/0xe0 [xfs]
[<f3cefa29>] xfs_trans_read_buf_map+0x179/0x3e0 [xfs]
[<f3cde727>] xfs_imap_to_bp+0x67/0xe0 [xfs]
[<f3cdec7d>] xfs_iread+0x7d/0x3e0 [xfs]
[<f3c996e8>] ? xfs_inode_alloc+0x58/0x1b0 [xfs]
[<f3c9a092>] xfs_iget+0x192/0x580 [xfs]
[<f3caa219>] ? kmem_free+0x19/0x50 [xfs]
[<f3ca0ce6>] xfs_bulkstat_one_int+0x86/0x2c0 [xfs]
[<f3ca0f54>] xfs_bulkstat_one+0x34/0x40 [xfs]
[<f3ca0c10>] ? xfs_internal_inum+0xa0/0xa0 [xfs]
[<f3ca1360>] xfs_bulkstat+0x400/0x870 [xfs]
[<f3c9ae86>] xfs_ioc_bulkstat+0xb6/0x160 [xfs]
[<f3ca0f20>] ? xfs_bulkstat_one_int+0x2c0/0x2c0 [xfs]
[<f3c9ccb0>] ? xfs_ioc_swapext+0x160/0x160 [xfs]
[<f3c9d433>] xfs_file_ioctl+0x783/0xa60 [xfs]
[<c106bf9b>] ? __update_cpu_load+0xab/0xd0
[<c105dc58>] ? hrtimer_forward+0xa8/0x1b0
[<c12d4050>] ? timerqueue_add+0x50/0xb0
[<c108d143>] ? ktime_get+0x53/0xe0
[<c1094035>] ? clockevents_program_event+0x95/0x130
[<f3c9ccb0>] ? xfs_ioc_swapext+0x160/0x160 [xfs]
[<c11349b2>] do_vfs_ioctl+0x2e2/0x4c0
[<c1095699>] ? tick_program_event+0x29/0x30
[<c105e42c>] ? hrtimer_interrupt+0x13c/0x2a0
[<c12d8dfb>] ? lockref_put_or_lock+0xb/0x30
[<c1134c08>] SyS_ioctl+0x78/0x80
[<c1540fa8>] syscall_call+0x7/0x7
INFO: task md3_resync:24817 blocked for more than 120 seconds.
Tainted: P O 3.14.19fe-dirty #3
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
c2d95e08 00000046 00000092 00000000 00000020 00000092 1befb1af 000030e0
c17d6700 c17d6700 e9d3c6e0 d18a9950 c2d95ddc c12adbc5 00000000 00000001
00000000 00000001 c2d95dec e95f0000 00000001 c2d95e08 00000292 00000292
Call Trace:
[<c12adbc5>] ? queue_unplugged+0x45/0x90
[<c10798bb>] ? prepare_to_wait_event+0x6b/0xd0
[<c153e893>] schedule+0x23/0x60
[<c145782d>] md_do_sync+0x8dd/0x1010
[<c1079990>] ? __wake_up_sync+0x20/0x20
[<c1453fb0>] ? md_start_sync+0xb0/0xb0
[<c1454098>] md_thread+0xe8/0x100
[<c107945f>] ? __wake_up_locked+0x1f/0x30
[<c1453fb0>] ? md_start_sync+0xb0/0xb0
[<c105ae21>] kthread+0xa1/0xc0
[<c15418b7>] ret_from_kernel_thread+0x1b/0x28
[<c105ad80>] ? kthread_create_on_node+0x110/0x110
Again, the blocked task info is edited for space; hopefully these are
the important ones,
but I've got more if it helps.
Thanks,
Bill
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Roberto Spadim @ 2014-09-30 20:29 UTC (permalink / raw)
To: Robert L Mathews; +Cc: Linux-RAID
In-Reply-To: <542B0B66.9080506@tigertech.com>
wow, nice :)
i think that's something near what i have today, the latency is the
main problem that i need to reduce and control, checking every time
with users, they don't care if system is slow to send a big file, they
care if system stop for some seconds and get back, and stop again and
get back, a smothly operation is prefered instead of batchs of high
speed
+1 to this experience, here i have some "D" state process too,
interesting common point
let'me ask others questions....
today you have 3 disk and/or ssd, before you had 3 disks, any time you
got 2 disks setup and tried to change only one disk to ssd? that will
be my scenario if i stay with raid and ssd solution without new cache
layers, about trim commands and this setup life time, there's any
scheduler cleanup method? how many time running this setup? here 1
year is a nice time to consider as a sucessfull project, i replace
disks near to 2 years
2014-09-30 16:58 GMT-03:00 Robert L Mathews <lists@tigertech.com>:
> On 9/30/14 9:20 AM, Roberto Spadim wrote:
>
>> well, others experiences and comments are wellcome :)
>> comments about cache (bcache,flashcache,dm cache) are wellcome too
>
> Keep in mind that sequential read and write speeds are not everything.
> For our use pattern, for example (busy Web and mail servers with
> millions of files that aren't necessarily grouped physically on the
> disk, even when in the same directory [think maildirs where each file
> represents one message]), latency is more important -- particularly read
> latency.
>
> Our servers use three-disk RAID 1 arrays. When we replaced one of the
> spinning hard drives in each array with an SSD (some of which were
> Samsung 840 Pros), then marked the remaining two spinning disks as
> "write-mostly", the average read latency on the array dropped from
> around 12 ms to less than 2 ms.
>
> More importantly, the average amount of time any process on a server is
> waiting in the "D" state (read or write) dropped from 8% to 3%.
>
> Note that improving the read performance this way also improves the
> write performance of the entire array, because when a write occurs, it
> will never be queued behind a spinning disk read: the spinning disk is
> more likely to be idle.
>
> So our experience confirms that adding a single SSD to a RAID 1 array,
> then marking the others write-mostly, is a good stopgap measure on the
> road to replacing all spinning disks with SSDs. It effectively doubled
> the average performance of our storage. No additional caching layers
> required.
>
> --
> Robert L Mathews, Tiger Technologies, http://www.tigertech.net/
> --
> To unsubscribe from this list: send the line "unsubscribe linux-raid" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
--
Roberto Spadim
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Robert L Mathews @ 2014-09-30 19:58 UTC (permalink / raw)
To: Linux-RAID
In-Reply-To: <CAH3kUhFwakMSRAgKdSR-4g53KwHS6Yj_mmm3TOVKRQ_qaA9SNA@mail.gmail.com>
On 9/30/14 9:20 AM, Roberto Spadim wrote:
> well, others experiences and comments are wellcome :)
> comments about cache (bcache,flashcache,dm cache) are wellcome too
Keep in mind that sequential read and write speeds are not everything.
For our use pattern, for example (busy Web and mail servers with
millions of files that aren't necessarily grouped physically on the
disk, even when in the same directory [think maildirs where each file
represents one message]), latency is more important -- particularly read
latency.
Our servers use three-disk RAID 1 arrays. When we replaced one of the
spinning hard drives in each array with an SSD (some of which were
Samsung 840 Pros), then marked the remaining two spinning disks as
"write-mostly", the average read latency on the array dropped from
around 12 ms to less than 2 ms.
More importantly, the average amount of time any process on a server is
waiting in the "D" state (read or write) dropped from 8% to 3%.
Note that improving the read performance this way also improves the
write performance of the entire array, because when a write occurs, it
will never be queued behind a spinning disk read: the spinning disk is
more likely to be idle.
So our experience confirms that adding a single SSD to a RAID 1 array,
then marking the others write-mostly, is a good stopgap measure on the
road to replacing all spinning disks with SSDs. It effectively doubled
the average performance of our storage. No additional caching layers
required.
--
Robert L Mathews, Tiger Technologies, http://www.tigertech.net/
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Roberto Spadim @ 2014-09-30 16:20 UTC (permalink / raw)
To: Andrei Banu; +Cc: Linux-RAID
In-Reply-To: <542AD06C.9080908@redhost.ro>
nice :) well today it use hdd with sata, 500GB with raid 1 (2 disks)
the only problem is the disk being old, i must replace it
i'm considering ssd only by the price and myth, hdd works well today
i have some doubts about performace without the raid card as you told,
the last ssd i used was enterprise level not desktop/non enterprise
level, and with a raid card
well, others experiences and comments are wellcome :)
comments about cache (bcache,flashcache,dm cache) are wellcome too
2014-09-30 12:46 GMT-03:00 Andrei Banu <andrei.banu@redhost.ro>:
> Hi,
>
> I am in no way of the same technical caliber as many of the people on this
> list so please take my advice with a grain of salt. However I do have some
> experience with Samsung SSDs in RAID 1 software (md-raid) and I would like
> to warn you against it.
>
> My setup is with 840 PROs of 512GB. I did leave free space (not partitioned)
> for OP. The SSDs are connected on SATA3.
>
> I'll give you the result of 2 tests and if you are happy, go ahead:
>
> READ
> root [~]# hdparm -t /dev/sda
> /dev/sda:
> Timing buffered disk reads: 292 MB in 3.01 seconds = 96.92 MB/sec
>
> WRITE
> root [~]# dd if=/dev/zero of=testn bs=4k count=256k conv=fdatasync
> 262144+0 records in
> 262144+0 records out
> 1073741824 bytes (1.1 GB) copied, 13.4047 s, 80.1 MB/s
>
> When the setup is fresh you'll get significantly more out of it but after a
> while
> (this setup is roughly 1 year old) this is what you get with Samsung SSDs in
> RAID1 SW. At least this is my experience and I did try to improve it. As a
> matter
> of fact, before making a secure erase of one of the SSDs the write speed was
> under 10MB/s. What you see above is a lot better than what it started out
> (but it's true that after the secure erase the read/write speeds were a lot
> better but it dropped in a very short time frame).
>
> A few points:
> 1. CAUTION: 840 EVO are proved to have a read speed degradation for old
> data written to the drive (this is unrelated to RAID but it probably affects
> performance in RAID).
> 2. I believe that any SSDs, regardless of the brand, in a software RAID-1
> setup
> might be a bad idea.
> 3. I believe that the same SSDs in a hardware RAID setup might lead to a
> different story.
>
> Again: I am not a technical wiz (like many of the people on this list) but I
> did
> have some experience with this so I thought I should let you know.
>
> Kind regards!
>
>
>
>
> On 30.09.2014 17:50, Roberto Spadim wrote:
>>
>> hi guys!
>> i will use a ssd raid1, i want know if raid1 trim is supported at mdadm
>> i will use a 840 evo (or evo pro not selected the right one yet) 500gb
>> each, raid1, today database size is 100gb, i think it will grow
>> 10gb/year, i had many space...
>>
>> the point are: madm raid1 trim is supported? or should i use lvm?
>> should i partition it with 400gb and leave 100gb untouched? or should
>> i use a hdd+ssd and dmcache?
>>
>>
>> :) thanks guys, that's a small enterprise solution, they can't buy
>> raid cards and sas harddisk are same price of ssd :)
>>
>> any idea/experience and information is wellcome
>>
>>
>
--
Roberto Spadim
SPAEmpresarial
Eng. Automação e Controle
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
^ permalink raw reply
* Re: raid1 - ssd, doubts
From: Andrei Banu @ 2014-09-30 15:46 UTC (permalink / raw)
To: Roberto Spadim, Linux-RAID
In-Reply-To: <CAH3kUhHkot3doqVZQ0TAot0A3e8i8RFWqsiLL35p1OYsvz5Xkw@mail.gmail.com>
Hi,
I am in no way of the same technical caliber as many of the people on this
list so please take my advice with a grain of salt. However I do have some
experience with Samsung SSDs in RAID 1 software (md-raid) and I would like
to warn you against it.
My setup is with 840 PROs of 512GB. I did leave free space (not partitioned)
for OP. The SSDs are connected on SATA3.
I'll give you the result of 2 tests and if you are happy, go ahead:
READ
root [~]# hdparm -t /dev/sda
/dev/sda:
Timing buffered disk reads: 292 MB in 3.01 seconds = 96.92 MB/sec
WRITE
root [~]# dd if=/dev/zero of=testn bs=4k count=256k conv=fdatasync
262144+0 records in
262144+0 records out
1073741824 bytes (1.1 GB) copied, 13.4047 s, 80.1 MB/s
When the setup is fresh you'll get significantly more out of it but
after a while
(this setup is roughly 1 year old) this is what you get with Samsung SSDs in
RAID1 SW. At least this is my experience and I did try to improve it. As
a matter
of fact, before making a secure erase of one of the SSDs the write speed
was
under 10MB/s. What you see above is a lot better than what it started out
(but it's true that after the secure erase the read/write speeds were a lot
better but it dropped in a very short time frame).
A few points:
1. CAUTION: 840 EVO are proved to have a read speed degradation for old
data written to the drive (this is unrelated to RAID but it probably
affects
performance in RAID).
2. I believe that any SSDs, regardless of the brand, in a software
RAID-1 setup
might be a bad idea.
3. I believe that the same SSDs in a hardware RAID setup might lead to a
different story.
Again: I am not a technical wiz (like many of the people on this list)
but I did
have some experience with this so I thought I should let you know.
Kind regards!
On 30.09.2014 17:50, Roberto Spadim wrote:
> hi guys!
> i will use a ssd raid1, i want know if raid1 trim is supported at mdadm
> i will use a 840 evo (or evo pro not selected the right one yet) 500gb
> each, raid1, today database size is 100gb, i think it will grow
> 10gb/year, i had many space...
>
> the point are: madm raid1 trim is supported? or should i use lvm?
> should i partition it with 400gb and leave 100gb untouched? or should
> i use a hdd+ssd and dmcache?
>
>
> :) thanks guys, that's a small enterprise solution, they can't buy
> raid cards and sas harddisk are same price of ssd :)
>
> any idea/experience and information is wellcome
>
>
^ permalink raw reply
* raid1 - ssd, doubts
From: Roberto Spadim @ 2014-09-30 14:50 UTC (permalink / raw)
To: Linux-RAID
hi guys!
i will use a ssd raid1, i want know if raid1 trim is supported at mdadm
i will use a 840 evo (or evo pro not selected the right one yet) 500gb
each, raid1, today database size is 100gb, i think it will grow
10gb/year, i had many space...
the point are: madm raid1 trim is supported? or should i use lvm?
should i partition it with 400gb and leave 100gb untouched? or should
i use a hdd+ssd and dmcache?
:) thanks guys, that's a small enterprise solution, they can't buy
raid cards and sas harddisk are same price of ssd :)
any idea/experience and information is wellcome
--
Roberto Spadim
^ permalink raw reply
* Re: /sys/block/md126 still exists even after stopping the array
From: Francis Moreau @ 2014-09-30 7:43 UTC (permalink / raw)
To: NeilBrown; +Cc: linux-raid, sebastian.riemer
In-Reply-To: <20140930075643.34e864fa@notabene.brown>
Hi Neil,
On 09/29/2014 11:56 PM, NeilBrown wrote:
> On Mon, 29 Sep 2014 10:45:17 +0200 Francis Moreau <francis.moro@gmail.com>
> wrote:
>
>>> So what were pids 930 and 459?
>>> One was presumably the "mdadm -Ss" - probably 930.
>>> Is 459 the "mdadm --monitor" ?? That might be useful hint.
>>>
>>
>> yes.
>>
>> [456] is: /sbin/mdadm --monitor --scan --daemonise --syslog
>> --pid-file=/run/mdadm/mdadm.pid
>>
>> and [930] is 'mdamd -Ss'.
>
> Good. Please try the patch below.
>
After applying your patch, this is what I'm getting in syslog:
Sep 30 03:40:07 localhost kernel: md_open(): md125 opened by mdadm [970]
Sep 30 03:40:07 localhost kernel: md_release(): md125 released by mdadm
[970]
Sep 30 03:40:07 localhost kernel: md_open(): md125 opened by mdadm [972]
Sep 30 03:40:07 localhost kernel: md_open(): md125 opened by mdadm [970]
Sep 30 03:40:07 localhost kernel: md_release(): md125 released by mdadm
[972]
Sep 30 03:40:07 localhost kernel: md_open(): md125 opened by
systemd-udevd [971]
Sep 30 03:40:07 localhost systemd[1]: Cannot add dependency job for unit
mdmonitor-takeover.service, ignoring: Invalid argument
Sep 30 03:40:07 localhost systemd[1]: Started Software RAID monitoring
and management.
Sep 30 03:40:07 localhost kernel: md_release(): md125 released by
systemd-udevd [971]
Sep 30 03:40:08 localhost mdadm[466]: DeviceDisappeared event detected
on md device /dev/md125
Sep 30 03:40:08 localhost mdadm[466]: DeviceDisappeared event detected
on md device /dev/md126
Sep 30 03:40:08 localhost mdadm[466]: DeviceDisappeared event detected
on md device /dev/md127
Sep 30 03:40:08 localhost kernel: md125: detected capacity change from
1863254016 to 0
Sep 30 03:40:08 localhost kernel: md: md125 stopped.
Sep 30 03:40:08 localhost kernel: md: unbind<vdc3>
Sep 30 03:40:08 localhost kernel: md: export_rdev(vdc3)
Sep 30 03:40:08 localhost kernel: md: unbind<vdb3>
Sep 30 03:40:08 localhost kernel: md: export_rdev(vdb3)
Sep 30 03:40:08 localhost kernel: md_release(): md125 released by mdadm
[970]
Sep 30 03:40:08 localhost kernel: md_open(): md127 opened by mdadm [466]
Sep 30 03:40:08 localhost kernel: md_release(): md127 released by mdadm
[466]
Sep 30 03:40:08 localhost kernel: md_open(): md126 opened by mdadm [466]
Sep 30 03:40:08 localhost kernel: md_release(): md126 released by mdadm
[466]
Sep 30 03:40:08 localhost kernel: md_open(): md126 opened by mdadm [970]
Sep 30 03:40:08 localhost kernel: md_release(): md126 released by mdadm
[970]
Sep 30 03:40:08 localhost kernel: md_open(): md126 opened by mdadm [970]
Sep 30 03:40:08 localhost kernel: md126: detected capacity change from
67043328 to 0
Sep 30 03:40:08 localhost kernel: md: md126 stopped.
Sep 30 03:40:08 localhost kernel: md: unbind<vdc1>
Sep 30 03:40:08 localhost kernel: md: export_rdev(vdc1)
Sep 30 03:40:08 localhost kernel: md: unbind<vdb1>
Sep 30 03:40:08 localhost kernel: md: export_rdev(vdb1)
Sep 30 03:40:08 localhost kernel: md_open(): md127 opened by mdadm [466]
Sep 30 03:40:08 localhost kernel: md_release(): md127 released by mdadm
[466]
Sep 30 03:40:08 localhost kernel: md_release(): md126 released by mdadm
[970]
Sep 30 03:40:08 localhost kernel: md_open(): md127 opened by mdadm [970]
Sep 30 03:40:08 localhost kernel: md_release(): md127 released by mdadm
[970]
Sep 30 03:40:08 localhost kernel: md_open(): md127 opened by mdadm [970]
Sep 30 03:40:08 localhost kernel: md127: detected capacity change from
214564864 to 0
Sep 30 03:40:08 localhost kernel: md: md127 stopped.
Sep 30 03:40:08 localhost kernel: md: unbind<vdc2>
Sep 30 03:40:08 localhost kernel: md: export_rdev(vdc2)
Sep 30 03:40:08 localhost kernel: md: unbind<vdb2>
Sep 30 03:40:08 localhost kernel: md: export_rdev(vdb2)
Sep 30 03:40:08 localhost kernel: md_release(): md127 released by mdadm
[970]
The ghost device is no more present so your patch seems to have fixed my
issue. But I must admit I don't really understand what's going on :-/
Thanks
^ permalink raw reply
* Re: Raid5 hang in 3.14.19
From: BillStuff @ 2014-09-30 4:19 UTC (permalink / raw)
To: NeilBrown; +Cc: linux-raid
In-Reply-To: <20140930075950.1d1e3865@notabene.brown>
On 09/29/2014 04:59 PM, NeilBrown wrote:
> On Sun, 28 Sep 2014 23:28:17 -0500 BillStuff <billstuff2001@sbcglobal.net>
> wrote:
>
>> On 09/28/2014 11:08 PM, NeilBrown wrote:
>>> On Sun, 28 Sep 2014 22:56:19 -0500 BillStuff <billstuff2001@sbcglobal.net>
>>> wrote:
>>>
>>>> On 09/28/2014 09:25 PM, NeilBrown wrote:
>>>>> On Fri, 26 Sep 2014 17:33:58 -0500 BillStuff <billstuff2001@sbcglobal.net>
>>>>> wrote:
>>>>>
>>>>>> Hi Neil,
>>>>>>
>>>>>> I found something that looks similar to the problem described in
>>>>>> "Re: seems like a deadlock in workqueue when md do a flush" from Sept 14th.
>>>>>>
>>>>>> It's on 3.14.19 with 7 recent patches for fixing raid1 recovery hangs.
>>>>>>
>>>>>> on this array:
>>>>>> md3 : active raid5 sdf1[5] sde1[4] sdd1[3] sdc1[2] sdb1[1] sda1[0]
>>>>>> 104171200 blocks level 5, 64k chunk, algorithm 2 [6/6] [UUUUUU]
>>>>>> bitmap: 1/5 pages [4KB], 2048KB chunk
>>>>>>
>>>>>> I was running a test doing parallel kernel builds, read/write loops, and
>>>>>> disk add / remove / check loops,
>>>>>> on both this array and a raid1 array.
>>>>>>
>>>>>> I was trying to stress test your recent raid1 fixes, which went well,
>>>>>> but then after 5 days,
>>>>>> the raid5 array hung up with this in dmesg:
>>>>> I think this is different to the workqueue problem you mentioned, though as I
>>>>> don't know exactly what caused either I cannot be certain.
>>>>>
>>>>> From the data you provided it looks like everything is waiting on
>>>>> get_active_stripe(), or on a process that is waiting on that.
>>>>> That seems pretty common whenever anything goes wrong in raid5 :-(
>>>>>
>>>>> The md3_raid5 task is listed as blocked, but not stack trace is given.
>>>>> If the machine is still in the state, then
>>>>>
>>>>> cat /proc/1698/stack
>>>>>
>>>>> might be useful.
>>>>> (echo t > /proc/sysrq-trigger is always a good idea)
>>>> Might this help? I believe the array was doing a "check" when things
>>>> hung up.
>>> It looks like it was trying to start doing a 'check'.
>>> The 'resync' thread hadn't been started yet.
>>> What is 'kthreadd' doing?
>>> My guess is that it is in try_to_free_pages() waiting for writeout
>>> for some xfs file page onto the md array ... which won't progress until
>>> the thread gets started.
>>>
>>> That would suggest that we need an async way to start threads...
>>>
>>> Thanks,
>>> NeilBrown
>>>
>> I suspect your guess is correct:
> Thanks for the confirmation.
>
> I'm thinking of something like that. Very basic suggestion suggests it
> instantly crash.
>
> If you were to apply this patch and run your test for a week or two, that
> would increase my confidence (though of course testing doesn't prove the
> absence of bugs....)
>
> Thanks,
> NeilBrown
Got it running. I'll let you know if anything interesting happens.
Thanks,
Bill
>
>
> diff --git a/drivers/md/md.c b/drivers/md/md.c
> index a79e51d15c2b..580d4b97696c 100644
> --- a/drivers/md/md.c
> +++ b/drivers/md/md.c
> @@ -7770,6 +7770,33 @@ no_add:
> return spares;
> }
>
> +static void md_start_sync(struct work_struct *ws)
> +{
> + struct mddev *mddev = container_of(ws, struct mddev, del_work);
> +
> + mddev->sync_thread = md_register_thread(md_do_sync,
> + mddev,
> + "resync");
> + if (!mddev->sync_thread) {
> + printk(KERN_ERR "%s: could not start resync"
> + " thread...\n",
> + mdname(mddev));
> + /* leave the spares where they are, it shouldn't hurt */
> + clear_bit(MD_RECOVERY_SYNC, &mddev->recovery);
> + clear_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
> + clear_bit(MD_RECOVERY_REQUESTED, &mddev->recovery);
> + clear_bit(MD_RECOVERY_CHECK, &mddev->recovery);
> + clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
> + if (test_and_clear_bit(MD_RECOVERY_RECOVER,
> + &mddev->recovery))
> + if (mddev->sysfs_action)
> + sysfs_notify_dirent_safe(mddev->sysfs_action);
> + } else
> + md_wakeup_thread(mddev->sync_thread);
> + sysfs_notify_dirent_safe(mddev->sysfs_action);
> + md_new_event(mddev);
> +}
> +
> /*
> * This routine is regularly called by all per-raid-array threads to
> * deal with generic issues like resync and super-block update.
> @@ -7823,6 +7850,7 @@ void md_check_recovery(struct mddev *mddev)
>
> if (mddev_trylock(mddev)) {
> int spares = 0;
> + bool sync_starting = false;
>
> if (mddev->ro) {
> /* On a read-only array we can:
> @@ -7921,28 +7949,14 @@ void md_check_recovery(struct mddev *mddev)
> */
> bitmap_write_all(mddev->bitmap);
> }
> - mddev->sync_thread = md_register_thread(md_do_sync,
> - mddev,
> - "resync");
> - if (!mddev->sync_thread) {
> - printk(KERN_ERR "%s: could not start resync"
> - " thread...\n",
> - mdname(mddev));
> - /* leave the spares where they are, it shouldn't hurt */
> - clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
> - clear_bit(MD_RECOVERY_SYNC, &mddev->recovery);
> - clear_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
> - clear_bit(MD_RECOVERY_REQUESTED, &mddev->recovery);
> - clear_bit(MD_RECOVERY_CHECK, &mddev->recovery);
> - } else
> - md_wakeup_thread(mddev->sync_thread);
> - sysfs_notify_dirent_safe(mddev->sysfs_action);
> - md_new_event(mddev);
> + INIT_WORK(&mddev->del_work, md_start_sync);
> + queue_work(md_misc_wq, &mddev->del_work);
> + sync_starting = true;
> }
> unlock:
> wake_up(&mddev->sb_wait);
>
> - if (!mddev->sync_thread) {
> + if (!mddev->sync_thread && !sync_starting) {
> clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
> if (test_and_clear_bit(MD_RECOVERY_RECOVER,
> &mddev->recovery))
>
^ permalink raw reply
* Re: Raid5 hang in 3.14.19
From: NeilBrown @ 2014-09-29 21:59 UTC (permalink / raw)
To: BillStuff; +Cc: linux-raid
In-Reply-To: <5428DFE1.9080600@sbcglobal.net>
[-- Attachment #1: Type: text/plain, Size: 5497 bytes --]
On Sun, 28 Sep 2014 23:28:17 -0500 BillStuff <billstuff2001@sbcglobal.net>
wrote:
> On 09/28/2014 11:08 PM, NeilBrown wrote:
> > On Sun, 28 Sep 2014 22:56:19 -0500 BillStuff <billstuff2001@sbcglobal.net>
> > wrote:
> >
> >> On 09/28/2014 09:25 PM, NeilBrown wrote:
> >>> On Fri, 26 Sep 2014 17:33:58 -0500 BillStuff <billstuff2001@sbcglobal.net>
> >>> wrote:
> >>>
> >>>> Hi Neil,
> >>>>
> >>>> I found something that looks similar to the problem described in
> >>>> "Re: seems like a deadlock in workqueue when md do a flush" from Sept 14th.
> >>>>
> >>>> It's on 3.14.19 with 7 recent patches for fixing raid1 recovery hangs.
> >>>>
> >>>> on this array:
> >>>> md3 : active raid5 sdf1[5] sde1[4] sdd1[3] sdc1[2] sdb1[1] sda1[0]
> >>>> 104171200 blocks level 5, 64k chunk, algorithm 2 [6/6] [UUUUUU]
> >>>> bitmap: 1/5 pages [4KB], 2048KB chunk
> >>>>
> >>>> I was running a test doing parallel kernel builds, read/write loops, and
> >>>> disk add / remove / check loops,
> >>>> on both this array and a raid1 array.
> >>>>
> >>>> I was trying to stress test your recent raid1 fixes, which went well,
> >>>> but then after 5 days,
> >>>> the raid5 array hung up with this in dmesg:
> >>> I think this is different to the workqueue problem you mentioned, though as I
> >>> don't know exactly what caused either I cannot be certain.
> >>>
> >>> From the data you provided it looks like everything is waiting on
> >>> get_active_stripe(), or on a process that is waiting on that.
> >>> That seems pretty common whenever anything goes wrong in raid5 :-(
> >>>
> >>> The md3_raid5 task is listed as blocked, but not stack trace is given.
> >>> If the machine is still in the state, then
> >>>
> >>> cat /proc/1698/stack
> >>>
> >>> might be useful.
> >>> (echo t > /proc/sysrq-trigger is always a good idea)
> >> Might this help? I believe the array was doing a "check" when things
> >> hung up.
> > It looks like it was trying to start doing a 'check'.
> > The 'resync' thread hadn't been started yet.
> > What is 'kthreadd' doing?
> > My guess is that it is in try_to_free_pages() waiting for writeout
> > for some xfs file page onto the md array ... which won't progress until
> > the thread gets started.
> >
> > That would suggest that we need an async way to start threads...
> >
> > Thanks,
> > NeilBrown
> >
>
> I suspect your guess is correct:
Thanks for the confirmation.
I'm thinking of something like that. Very basic suggestion suggests it
instantly crash.
If you were to apply this patch and run your test for a week or two, that
would increase my confidence (though of course testing doesn't prove the
absence of bugs....)
Thanks,
NeilBrown
diff --git a/drivers/md/md.c b/drivers/md/md.c
index a79e51d15c2b..580d4b97696c 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -7770,6 +7770,33 @@ no_add:
return spares;
}
+static void md_start_sync(struct work_struct *ws)
+{
+ struct mddev *mddev = container_of(ws, struct mddev, del_work);
+
+ mddev->sync_thread = md_register_thread(md_do_sync,
+ mddev,
+ "resync");
+ if (!mddev->sync_thread) {
+ printk(KERN_ERR "%s: could not start resync"
+ " thread...\n",
+ mdname(mddev));
+ /* leave the spares where they are, it shouldn't hurt */
+ clear_bit(MD_RECOVERY_SYNC, &mddev->recovery);
+ clear_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
+ clear_bit(MD_RECOVERY_REQUESTED, &mddev->recovery);
+ clear_bit(MD_RECOVERY_CHECK, &mddev->recovery);
+ clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
+ if (test_and_clear_bit(MD_RECOVERY_RECOVER,
+ &mddev->recovery))
+ if (mddev->sysfs_action)
+ sysfs_notify_dirent_safe(mddev->sysfs_action);
+ } else
+ md_wakeup_thread(mddev->sync_thread);
+ sysfs_notify_dirent_safe(mddev->sysfs_action);
+ md_new_event(mddev);
+}
+
/*
* This routine is regularly called by all per-raid-array threads to
* deal with generic issues like resync and super-block update.
@@ -7823,6 +7850,7 @@ void md_check_recovery(struct mddev *mddev)
if (mddev_trylock(mddev)) {
int spares = 0;
+ bool sync_starting = false;
if (mddev->ro) {
/* On a read-only array we can:
@@ -7921,28 +7949,14 @@ void md_check_recovery(struct mddev *mddev)
*/
bitmap_write_all(mddev->bitmap);
}
- mddev->sync_thread = md_register_thread(md_do_sync,
- mddev,
- "resync");
- if (!mddev->sync_thread) {
- printk(KERN_ERR "%s: could not start resync"
- " thread...\n",
- mdname(mddev));
- /* leave the spares where they are, it shouldn't hurt */
- clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
- clear_bit(MD_RECOVERY_SYNC, &mddev->recovery);
- clear_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
- clear_bit(MD_RECOVERY_REQUESTED, &mddev->recovery);
- clear_bit(MD_RECOVERY_CHECK, &mddev->recovery);
- } else
- md_wakeup_thread(mddev->sync_thread);
- sysfs_notify_dirent_safe(mddev->sysfs_action);
- md_new_event(mddev);
+ INIT_WORK(&mddev->del_work, md_start_sync);
+ queue_work(md_misc_wq, &mddev->del_work);
+ sync_starting = true;
}
unlock:
wake_up(&mddev->sb_wait);
- if (!mddev->sync_thread) {
+ if (!mddev->sync_thread && !sync_starting) {
clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
if (test_and_clear_bit(MD_RECOVERY_RECOVER,
&mddev->recovery))
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 828 bytes --]
^ permalink raw reply related
* Re: /sys/block/md126 still exists even after stopping the array
From: NeilBrown @ 2014-09-29 21:56 UTC (permalink / raw)
To: Francis Moreau; +Cc: linux-raid, sebastian.riemer
In-Reply-To: <54291C1D.7010005@gmail.com>
[-- Attachment #1: Type: text/plain, Size: 3512 bytes --]
On Mon, 29 Sep 2014 10:45:17 +0200 Francis Moreau <francis.moro@gmail.com>
wrote:
> > So what were pids 930 and 459?
> > One was presumably the "mdadm -Ss" - probably 930.
> > Is 459 the "mdadm --monitor" ?? That might be useful hint.
> >
>
> yes.
>
> [456] is: /sbin/mdadm --monitor --scan --daemonise --syslog
> --pid-file=/run/mdadm/mdadm.pid
>
> and [930] is 'mdamd -Ss'.
Good. Please try the patch below.
> >
> >>
> >>
> >>> Probably there is a 'change' event happening just before the 'remove' event,
> >>> and udev runs "mdadm" on the 'change' event, and that ends up happening after
> >>> the device has been removed.
> >>>
> >>> Is this really a problem? Can't you just ignore it and pretend it isn't
> >>> there?
> >>
> >> Well, if you list the block devices that the kernel detected in order to
> >> operate on them, it could. I don't know exactly what would be the result
> >> to use it but it could confuse some tools.
> >>
> >> Is there a way to check that the 'ghost' device has been removed by
> >> poking sysfs ?
> >
> > If you look at /sys/block/md*/md/array_state, those that contain 'inactive'
> > or 'clear' might be 'ghosts', or might be in the process of being assembled.
> > If you write 'clear' to the same file they should disappear.... unless udev
> > does something to re-create them.
> >
>
> It's in 'clear' state, and writing 'clear' doesn't make the device disapear.
>
> [root@localhost ~]# dmesg -c >/dev/null
> [root@localhost ~]# echo clear >/sys/block/md125/md/array_state
> [root@localhost ~]# dmesg
> [ 254.106252] md: md125 stopped.
> [ 254.108182] md_open(): mdX opened by mdadm [968]
>
> [ 254.109103] md_open(): md125 opened by mdadm [459]
> [ 254.109127] md_open(): md125 opened by mdadm [459]
> [ 254.109281] md_release(): md125 released by mdadm [459]
>
> [ 254.109337] md_open(): md125 opened by mdadm [968]
> [ 254.109572] md_release(): md125 released by mdadm [968]
>
> [ 254.109847] md_open(): md125 opened by systemd-udevd [967]
> [ 254.109986] md_release(): md125 released by systemd-udevd [967]
>
> In that sequence, it seems that mdadm [459] is missing a md_release()
> here. Is this expected ?
Presumably the first md_open returned an error. You could add another printk
at each 'return' to check.
Thanks,
NeilBrown
diff --git a/Monitor.c b/Monitor.c
index 5cb24fab8f2a..971d2ecbea72 100644
--- a/Monitor.c
+++ b/Monitor.c
@@ -460,7 +460,7 @@ static int check_array(struct state *st, struct mdstat_ent *mdstat,
mdu_array_info_t array;
struct mdstat_ent *mse = NULL, *mse2;
char *dev = st->devname;
- int fd;
+ int fd = -1;
int i;
int remaining_disks;
int last_disk;
@@ -468,6 +468,27 @@ static int check_array(struct state *st, struct mdstat_ent *mdstat,
if (test)
alert("TestMessage", dev, NULL, ainfo);
+ if (st->devnm[0])
+ fd = open("/sys/block", O_RDONLY|O_DIRECTORY);
+ if (fd >= 0) {
+ /* Don't open the device unless it is present and
+ * active in sysfs.
+ */
+ char buf[10];
+ close(fd);
+ fd = sysfs_open(st->devnm, NULL, "array_state");
+ if (fd < 0 ||
+ read(fd, buf, 10) < 5 ||
+ strncmp(buf,"clear",5) == 0 ||
+ strncmp(buf,"inact",5) == 0) {
+ if (fd >= 0)
+ close(fd);
+ if (!st->err)
+ alert("DeviceDisappeared", dev, NULL, ainfo);
+ st->err++;
+ return 0;
+ }
+ }
fd = open(dev, O_RDONLY);
if (fd < 0) {
if (!st->err)
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 828 bytes --]
^ permalink raw reply related
* Re: dmadm question
From: NeilBrown @ 2014-09-29 21:30 UTC (permalink / raw)
To: Dan Williams; +Cc: luke, linux-raid, Artur Paszkiewicz
In-Reply-To: <CAA9_cmc_YXg1=uVCm1wQ999ZMOihEv6HB-AO0FL8aiTwvFCQag@mail.gmail.com>
[-- Attachment #1: Type: text/plain, Size: 648 bytes --]
On Mon, 29 Sep 2014 10:38:03 -0700 Dan Williams <dan.j.williams@intel.com>
wrote:
> Hmm, I wonder if this is a regression. The last time I saw this
> problem I introduced commit:
>
Maybe...
Luke, can you change get_extents() in super-intel.c like this:
if (dl->index == -1)
+ reservation = MPB_SECTOR_CNT;
- reservation = imsm_min_reserved_sectors(super);
else
reservation = MPB_SECTOR_CNT + IMSM_RESERVED_SECTORS;
i.e. replace imsm_min_reserved_sectors(super) with MPB_SECTOR_CNT
and 'make install'.
Then see if your problem goes away?
Thanks,
NeilBrown
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 828 bytes --]
^ permalink raw reply
* Re: dmadm question
From: Dan Williams @ 2014-09-29 17:38 UTC (permalink / raw)
To: luke; +Cc: NeilBrown, linux-raid, Artur Paszkiewicz
In-Reply-To: <56516488c20b6b78729f14731fe8aecf.squirrel@webmail.lukeodom.com>
Hmm, I wonder if this is a regression. The last time I saw this
problem I introduced commit:
commit b276dd33c74a51598e37fc72e6fb8f5ebd6620f2
Author: Dan Williams <dan.j.williams@intel.com>
Date: Thu Aug 25 19:14:24 2011 -0700
imsm: fix reserved sectors for spares
Different OROMs reserve different amounts of space for the migration area.
When activating a spare minimize the reserved space otherwise a valid spare
can be prevented from joining an array with a migration area smaller than
IMSM_RESERVED_SECTORS.
This may result in an array that cannot be reshaped, but that is less
surprising than not being able to rebuild a degraded array.
imsm_reserved_sectors() already reports the minimal value which adds to
the confusion when trying rebuild an array because mdadm -E indicates
that the device has enough space.
Cc: Anna Czarnowska <anna.czarnowska@intel.com>
Signed-off-by: Dan Williams <dan.j.williams@intel.com>
Signed-off-by: NeilBrown <neilb@suse.de>
...but I notice now the logic was replaced with:
commit b81221b74eba9fd7f670a8d3d4bfbda43ec91993
Author: Czarnowska, Anna <anna.czarnowska@intel.com>
Date: Mon Sep 19 12:57:48 2011 +0000
imsm: Calculate reservation for a spare based on active disks in container
New function to calculate minimum reservation to expect from a spare
is introduced.
The required amount of space at the end of the disk depends on what we
plan to do with the spare and what array we want to use it in.
For creating new subarray in an empty container the full reservation of
MPB_SECTOR_COUNT + IMSM_RESERVED_SECTORS is required.
For recovery or OLCE on a volume using new metadata format at least
MPB_SECTOR_CNT + NUM_BLOCKS_DIRTY_STRIPE_REGION is required.
The additional space for migration optimization included in
IMSM_RESERVED_SECTORS is not necessary and is not reserved by some oroms.
MPB_SECTOR_CNT alone is not sufficient as it does not include the
reservation at the end of subarray.
However if the real reservation on active disks is smaller than this
(when the array uses old metadata format) we should use the real value.
This will allow OLCE and recovery to start on the spare even if the volume
doesn't have the reservation we normally use for new volumes.
Signed-off-by: Anna Czarnowska <anna.czarnowska@intel.com>
Signed-off-by: NeilBrown <neilb@suse.de>
...which raises the absolute minimum reserved sectors to a number >
MPB_SECTOR_CNT.
My thought is that we should always favor spare replacement over
features like reshape, and dirty-stripe-log. Commit b81221b74eba
appears to favor the latter.
On Mon, Sep 15, 2014 at 7:07 AM, Luke Odom <luke@lukeodom.com> wrote:
> Drive is exact same model as old one. Output of requested commands:
>
> # mdadm --manage /dev/md127 --remove /dev/sdb
> mdadm: hot removed /dev/sdb from /dev/md127
> # mdadm --zero /dev/sdb
> # mdadm --manage /dev/md127 --add /dev/sdb
> mdadm: added /dev/sdb
> # ps aux | grep mdmon
> root 1937 0.0 0.1 10492 10484 ? SLsl 14:04 0:00 mdmon md127
> root 2055 0.0 0.0 2420 928 pts/0 S+ 14:06 0:00 grep mdmon
>
> md: unbind<sdb>
> md: export_rdev(sdb)
> md: bind<sdb>
>
>
>
> On Sun, September 14, 2014 5:31 pm, NeilBrown wrote:
>> On 12 Sep 2014 18:49:54 -0700 Luke Odom <luke@lukeodom.com> wrote:
>>
>>> I had a raid1 subarray running within an imsm container. One of the
>>> drives died so I replaced it. I can get the new drive into the imsm
>>> container but I can’t add it to the raid1 array within that
>>> container. I’ve read the man page and can’t see to figure it out.
>>> Any help would be greatly appreciated. Using mdadm 3.2.5 on debian
>>> squeeze.
>>
>> This should just happen automatically. As soon as you add the device to
>> the
>> container, mdmon notices and adds it to the raid1.
>>
>> However it appears not to have happened...
>>
>> I assume the new drive is exactly the same size as the old drive?
>> Try removing the new device from md127, run "mdadm --zero" on it, then add
>> it
>> back again.
>> Do any messages appear in the kernel logs when you do that?
>>
>> Is "mdmon md127" running?
>>
>> NeilBrown
>>
>>
>>>
>>>
>>>
>>>
>>> root@ds6790:~# cat /proc/mdstat
>>> Personalities : [raid0] [raid1] [raid10] [raid6] [raid5] [raid4]
>>> md126 : active raid1 sda[0]
>>> 976759808 blocks super external:/md127/0 [2/1] [U_]
>>>
>>>
>>>
>>>
>>> md127 : inactive sdb[0](S) sda[1](S)
>>> 4901 blocks super external:imsm
>>>
>>>
>>>
>>>
>>> unused devices: <none>
>>>
>>>
>>>
>>>
>>>
>>>
>>> root@ds6790:~# mdadm --detail /dev/md126
>>> /dev/md126:
>>> Container : /dev/md127, member 0
>>> Raid Level : raid1
>>> Array Size : 976759808 (931.51 GiB 1000.20 GB)
>>> Used Dev Size : 976759940 (931.51 GiB 1000.20 GB)
>>> Raid Devices : 2
>>> Total Devices : 1
>>>
>>>
>>>
>>>
>>> State : active, degraded
>>> Active Devices : 1
>>> Working Devices : 1
>>> Failed Devices : 0
>>> Spare Devices : 0
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>> UUID : 1be60edf:5c16b945:86434b6b:2714fddb
>>> Number Major Minor RaidDevice State
>>> 0 8 0 0 active sync
>>> /dev/sda
>>> 1 0 0 1 removed
>>>
>>>
>>>
>>>
>>>
>>>
>>> root@ds6790:~# mdadm --examine /dev/md127
>>> /dev/md127:
>>> Magic : Intel Raid ISM Cfg Sig.
>>> Version : 1.1.00
>>> Orig Family : 6e37aa48
>>> Family : 6e37aa48
>>> Generation : 00640a43
>>> Attributes : All supported
>>> UUID : ac27ba68:f8a3618d:3810d44f:25031c07
>>> Checksum : 513ef1f6 correct
>>> MPB Sectors : 1
>>> Disks : 2
>>> RAID Devices : 1
>>>
>>>
>>>
>>>
>>> Disk00 Serial : 9XG3RTL0
>>> State : active
>>> Id : 00000002
>>> Usable Size : 1953519880 (931.51 GiB 1000.20 GB)
>>>
>>>
>>>
>>>
>>> [Volume0]:
>>> UUID : 1be60edf:5c16b945:86434b6b:2714fddb
>>> RAID Level : 1
>>> Members : 2
>>> Slots : [U_]
>>> Failed disk : 1
>>> This Slot : 0
>>> Array Size : 1953519616 (931.51 GiB 1000.20 GB)
>>> Per Dev Size : 1953519880 (931.51 GiB 1000.20 GB)
>>> Sector Offset : 0
>>> Num Stripes : 7630936
>>> Chunk Size : 64 KiB
>>> Reserved : 0
>>> Migrate State : idle
>>> Map State : degraded
>>> Dirty State : dirty
>>>
>>>
>>>
>>>
>>> Disk01 Serial : XG3RWMF
>>> State : failed
>>> Id : ffffffff
>>> Usable Size : 1953519880 (931.51 GiB 1000.20 GB)
>>>
>>>
>>>
>>>
>>>
>>
>>
>
> --
> To unsubscribe from this list: send the line "unsubscribe linux-raid" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
^ permalink raw reply
* Re: /sys/block/md126 still exists even after stopping the array
From: Francis Moreau @ 2014-09-29 8:45 UTC (permalink / raw)
To: NeilBrown; +Cc: linux-raid, sebastian.riemer
In-Reply-To: <20140929143735.5fa54253@notabene.brown>
Hello Neil,
On 09/29/2014 06:37 AM, NeilBrown wrote:
> On Fri, 26 Sep 2014 14:21:04 +0200 Francis Moreau <francis.moro@gmail.com>
> wrote:
>
>> On 09/26/2014 12:44 PM, NeilBrown wrote:
>>> On Fri, 26 Sep 2014 12:23:27 +0200 Francis Moreau <francis.moro@gmail.com>
>>> wrote:
>>>
>>>> Hello Neil,
>>>>
>>>> On 09/26/2014 02:33 AM, NeilBrown wrote:
>>>>> On Thu, 25 Sep 2014 18:12:07 +0200 Francis Moreau <francis.moro@gmail.com>
>>>>> wrote:
>>>> [...]
>>>>>> I tried to find out what could have opened the md device by using fuser,
>>>>>> but fuser reports no users.
>>>>>
>>>>> It is probably a transient open/close.
>>>>>
>>>>
>>>> If it's open/close wouldn't the 'close' part make the device disapear ?
>>>
>>> No. It's ... complicated.
>>>
>>>>
>>>>>>
>>>>>> I took a look to the udev rules which are the one shipped by mdadm 3.3.2
>>>>>> but nothing keep the device opened during the remove event.
>>>>>>
>>>>>> Could you give me some hints here to debug this ?
>>>>>
>>>>> Modify md_open in drivers/md/md.c to add
>>>>> printk("Opened by %s\n", current->comm);
>>>>>
>>>>> and build a new kernel. That will tell you the name of the process which
>>>>> opened the device.
>>>>>
>>>>
>>>> I did that I also added a trace in md_release() but strangely no trace
>>>> were outputed from there.
>>>
>>> Without seeing your patch I can't guess what it happening, but I am *certain*
>>> that md_release() would get called providing md_open didn't return an error.
>>
>> Here's the patch:
>>
>> diff --git a/drivers/md/md.c b/drivers/md/md.c
>> index 73aedcb..08ead8d 100644
>> --- a/drivers/md/md.c
>> +++ b/drivers/md/md.c
>> @@ -6703,6 +6703,8 @@ static int md_open(struct block_device *bdev,
>> fmode_t mode)
>> struct mddev *mddev = mddev_find(bdev->bd_dev);
>> int err;
>>
>> + printk("md_open(): opened by %s\n", current->comm);
>> +
>> if (!mddev)
>> return -ENODEV;
>>
>> @@ -6735,6 +6737,8 @@ static void md_release(struct gendisk *disk,
>> fmode_t mode)
>> {
>> struct mddev *mddev = disk->private_data;
>>
>> + printk("md_release(): released by %s\n", current->comm);
>> +
>> BUG_ON(!mddev);
>> atomic_dec(&mddev->openers);
>> mddev_put(mddev);
>>
>>>
>>> It might be helpful to print out the pid and the md device number too
>>> task_tgid_vnr(current)
>>> will give you the pid.
>>> mdname(mddev)
>>> give the name of the device.
>>>
>>
>> Here's the new trace, this time md_release() was called, so I probably
>> did something wrong the first time, sorry for that.
>>
>> [ 1.470744] md_open(): md127 opened by mdadm [388]
>> [ 1.485437] md_release(): md127 released by mdadm [388]
>> [ 1.486888] md_open(): md126 opened by mdadm [381]
>> [ 1.487468] md_release(): md126 released by mdadm [381]
>> [ 1.488646] md_open(): md125 opened by mdadm [383]
>> [ 1.489074] md_release(): md125 released by mdadm [383]
>> [ 1.490555] md_open(): md127 opened by mdadm [385]
>> [ 1.512556] md_release(): md127 released by mdadm [385]
>> [ 1.512582] md_open(): md127 opened by mdadm [385]
>> [ 1.512682] md_open(): md126 opened by mdadm [384]
>> [ 1.553414] md_release(): md126 released by mdadm [384]
>> [ 1.553442] md_open(): md126 opened by mdadm [384]
>> [ 1.553549] md_open(): md125 opened by mdadm [382]
>> [ 1.573263] md_release(): md125 released by mdadm [382]
>> [ 1.573288] md_open(): md125 opened by mdadm [382]
>> [ 1.601034] md_open(): md125 opened by mdadm [459]
>> [ 1.601041] md_release(): md125 released by mdadm [459]
>> [ 1.601065] md_open(): md126 opened by mdadm [459]
>> [ 1.601067] md_release(): md126 released by mdadm [459]
>> [ 1.601090] md_open(): md127 opened by mdadm [459]
>> [ 1.601092] md_release(): md127 released by mdadm [459]
>> [ 1.601130] md_open(): md127 opened by mdadm [459]
>> [ 1.601220] md_release(): md127 released by mdadm [459]
>> [ 1.601633] md_open(): md126 opened by mdadm [459]
>> [ 1.601661] md_release(): md126 released by mdadm [459]
>> [ 1.601673] md_open(): md125 opened by mdadm [459]
>> [ 1.601695] md_release(): md125 released by mdadm [459]
>> [ 1.606127] md_open(): md125 opened by mdadm [454]
>> [ 1.608682] md_open(): md126 opened by mdadm [453]
>> [ 1.609514] md_open(): md127 opened by mdadm [448]
>> [ 1.622512] md_release(): md126 released by mdadm [453]
>> [ 1.623028] md_release(): md127 released by mdadm [448]
>> [ 1.625288] md_open(): md126 opened by systemd-udevd [363]
>> [ 1.625391] md_release(): md125 released by mdadm [454]
>> [ 1.625619] md_open(): md127 opened by systemd-udevd [368]
>> [ 1.625737] md_open(): md125 opened by systemd-udevd [366]
>> [ 1.637137] md_release(): md125 released by systemd-udevd [366]
>> [ 1.643982] md_open(): md125 opened by mdadm [476]
>> [ 1.644071] md_release(): md127 released by systemd-udevd [368]
>> [ 1.647787] md_release(): md125 released by mdadm [382]
>> [ 1.648171] md_release(): md126 released by systemd-udevd [363]
>> [ 1.651629] md_open(): md126 opened by mdadm [479]
>> [ 1.656666] md_open(): md127 opened by mdadm [480]
>> [ 1.657771] md_release(): md125 released by mdadm [476]
>> [ 1.659312] md_open(): md125 opened by systemd-udevd [365]
>> [ 1.663193] md_release(): md127 released by mdadm [385]
>> [ 1.673669] md_release(): md125 released by systemd-udevd [365]
>> [ 1.685527] md_release(): md127 released by mdadm [480]
>> [ 1.685599] md_release(): md126 released by mdadm [479]
>> [ 1.686058] md_open(): md126 opened by systemd-udevd [366]
>> [ 1.686282] md_release(): md126 released by systemd-udevd [366]
>> [ 1.691024] md_open(): md127 opened by systemd-udevd [363]
>> [ 1.695415] md_release(): md126 released by mdadm [384]
>> [ 1.707163] md_release(): md127 released by systemd-udevd [363]
>>
>>>>> mdadm --stop --scan <<<
>>
>> [ 89.975162] md_open(): md125 opened by mdadm [930]
>> [ 89.975305] md_release(): md125 released by mdadm [930]
>> [ 89.977434] md_open(): md125 opened by mdadm [932]
>> [ 89.978813] md_open(): md125 opened by mdadm [930]
>> [ 89.979365] md_release(): md125 released by mdadm [932]
>> [ 89.979693] md_open(): md125 opened by systemd-udevd [931]
>> [ 89.985790] md_release(): md125 released by systemd-udevd [931]
>> [ 90.179911] md_release(): md125 released by mdadm [930]
>> [ 90.180168] md_open(): md127 opened by mdadm [459]
>> [ 90.180187] md_release(): md127 released by mdadm [459]
>> [ 90.180199] md_open(): md126 opened by mdadm [459]
>> [ 90.180205] md_release(): md126 released by mdadm [459]
>> [ 90.180556] md_open(): md126 opened by mdadm [930]
>> [ 90.180653] md_release(): md126 released by mdadm [930]
>> [ 90.180690] md_open(): md126 opened by mdadm [930]
>> [ 90.180758] md_open(): mdX opened by mdadm [459]
>> [ 90.180995] md_open(): md125 opened by mdadm [459]
>> [ 90.181056] md_release(): md125 released by mdadm [459]opened by mdadm [968]
>> [ 90.182717] md_open(): md127 opened by mdadm [459]
>> [ 90.182725] md_release(): md127 released by mdadm [459]
>> [ 90.182732] md_open(): md126 opened by mdadm [459]
>> [ 90.182761] md_release(): md126 released by mdadm [459]
>> [ 90.182770] md_open(): md125 opened by mdadm [459]
>> [ 90.182775] md_release(): md125 released by mdadm [459]
>> [ 90.182940] md_release(): md126 released by mdadm [930]
>> [ 90.183167] md_open(): md127 opened by mdadm [930]
>> [ 90.183257] md_release(): md127 released by mdadm [930]
>> [ 90.183288] md_open(): md127 opened by mdadm [930]
>> [ 90.183461] md_open(): md127 opened by mdadm [459]
>> [ 90.183488] md_release(): md127 released by mdadm [459]
>> [ 90.183499] md_open(): md125 opened by mdadm [459]
>> [ 90.183505] md_release(): md125 released by mdadm [459]
>> [ 90.183686] md_release(): md127 released by mdadm [930]
>
> So what were pids 930 and 459?
> One was presumably the "mdadm -Ss" - probably 930.
> Is 459 the "mdadm --monitor" ?? That might be useful hint.
>
yes.
[456] is: /sbin/mdadm --monitor --scan --daemonise --syslog
--pid-file=/run/mdadm/mdadm.pid
and [930] is 'mdamd -Ss'.
>
>>
>>
>>> Probably there is a 'change' event happening just before the 'remove' event,
>>> and udev runs "mdadm" on the 'change' event, and that ends up happening after
>>> the device has been removed.
>>>
>>> Is this really a problem? Can't you just ignore it and pretend it isn't
>>> there?
>>
>> Well, if you list the block devices that the kernel detected in order to
>> operate on them, it could. I don't know exactly what would be the result
>> to use it but it could confuse some tools.
>>
>> Is there a way to check that the 'ghost' device has been removed by
>> poking sysfs ?
>
> If you look at /sys/block/md*/md/array_state, those that contain 'inactive'
> or 'clear' might be 'ghosts', or might be in the process of being assembled.
> If you write 'clear' to the same file they should disappear.... unless udev
> does something to re-create them.
>
It's in 'clear' state, and writing 'clear' doesn't make the device disapear.
[root@localhost ~]# dmesg -c >/dev/null
[root@localhost ~]# echo clear >/sys/block/md125/md/array_state
[root@localhost ~]# dmesg
[ 254.106252] md: md125 stopped.
[ 254.108182] md_open(): mdX opened by mdadm [968]
[ 254.109103] md_open(): md125 opened by mdadm [459]
[ 254.109127] md_open(): md125 opened by mdadm [459]
[ 254.109281] md_release(): md125 released by mdadm [459]
[ 254.109337] md_open(): md125 opened by mdadm [968]
[ 254.109572] md_release(): md125 released by mdadm [968]
[ 254.109847] md_open(): md125 opened by systemd-udevd [967]
[ 254.109986] md_release(): md125 released by systemd-udevd [967]
In that sequence, it seems that mdadm [459] is missing a md_release()
here. Is this expected ?
Thanks
^ permalink raw reply
* color box, display box, corrugated box, color card, blister card, color sleeve, hang tag, label
From: Jinghao Printing - CHINA @ 2014-09-29 6:35 UTC (permalink / raw)
Hi, this is David Wu from Shanghai, China.
We are a printing company, we can print color box, corrugated box,
label, hang tag etc.
Please let me know if you need these.
I will send you the website then.
Best regards,
David Wu
^ permalink raw reply
* Re: /sys/block/md126 still exists even after stopping the array
From: NeilBrown @ 2014-09-29 4:47 UTC (permalink / raw)
To: Francis Moreau; +Cc: linux-raid, sebastian.riemer
In-Reply-To: <54256100.3090507@gmail.com>
[-- Attachment #1: Type: text/plain, Size: 3413 bytes --]
On Fri, 26 Sep 2014 14:50:08 +0200 Francis Moreau <francis.moro@gmail.com>
wrote:
> On 09/26/2014 02:21 PM, Francis Moreau wrote:
> [...]
>
> >
> >>>> mdadm --stop --scan <<<
> >
> > [ 89.975162] md_open(): md125 opened by mdadm [930]
> > [ 89.975305] md_release(): md125 released by mdadm [930]
> > [ 89.977434] md_open(): md125 opened by mdadm [932]
> > [ 89.978813] md_open(): md125 opened by mdadm [930]
> > [ 89.979365] md_release(): md125 released by mdadm [932]
> > [ 89.979693] md_open(): md125 opened by systemd-udevd [931]
> > [ 89.985790] md_release(): md125 released by systemd-udevd [931]
> > [ 90.179911] md_release(): md125 released by mdadm [930]
> > [ 90.180168] md_open(): md127 opened by mdadm [459]
> > [ 90.180187] md_release(): md127 released by mdadm [459]
> > [ 90.180199] md_open(): md126 opened by mdadm [459]
> > [ 90.180205] md_release(): md126 released by mdadm [459]
> > [ 90.180556] md_open(): md126 opened by mdadm [930]
> > [ 90.180653] md_release(): md126 released by mdadm [930]
> > [ 90.180690] md_open(): md126 opened by mdadm [930]
> > [ 90.180758] md_open(): mdX opened by mdadm [459]
>
> What is this 'mdX' device that mdadm operates on ?
>
> It also doesn't have a counterpart release() call.
'mdX' is the name used if mddev->gendisk is NULL.
In that case, md_open() will return an error (ERESTARTSYS).
As the 'open' failed, we wouldn't expect a matching close/release.
NeilBrown
>
>
> > [ 90.180995] md_open(): md125 opened by mdadm [459]
> > [ 90.181056] md_release(): md125 released by mdadm [459]
> > [ 90.182717] md_open(): md127 opened by mdadm [459]
> > [ 90.182725] md_release(): md127 released by mdadm [459]
> > [ 90.182732] md_open(): md126 opened by mdadm [459]
> > [ 90.182761] md_release(): md126 released by mdadm [459]
> > [ 90.182770] md_open(): md125 opened by mdadm [459]
> > [ 90.182775] md_release(): md125 released by mdadm [459]
> > [ 90.182940] md_release(): md126 released by mdadm [930]
> > [ 90.183167] md_open(): md127 opened by mdadm [930]
> > [ 90.183257] md_release(): md127 released by mdadm [930]
> > [ 90.183288] md_open(): md127 opened by mdadm [930]
> > [ 90.183461] md_open(): md127 opened by mdadm [459]
> > [ 90.183488] md_release(): md127 released by mdadm [459]
> > [ 90.183499] md_open(): md125 opened by mdadm [459]
> > [ 90.183505] md_release(): md125 released by mdadm [459]
> > [ 90.183686] md_release(): md127 released by mdadm [930]
> >
> >
> >> Probably there is a 'change' event happening just before the 'remove' event,
> >> and udev runs "mdadm" on the 'change' event, and that ends up happening after
> >> the device has been removed.
> >>
> >> Is this really a problem? Can't you just ignore it and pretend it isn't
> >> there?
> >
> > Well, if you list the block devices that the kernel detected in order to
> > operate on them, it could. I don't know exactly what would be the result
> > to use it but it could confuse some tools.
> >
> > Is there a way to check that the 'ghost' device has been removed by
> > poking sysfs ?
> >
> > Thanks
> >
>
> --
> To unsubscribe from this list: send the line "unsubscribe linux-raid" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 828 bytes --]
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox