Linux RAID subsystem development
 help / color / mirror / Atom feed
* You have received fax, document 00000961727
From: Interfax @ 2015-11-14 10:26 UTC (permalink / raw)
  To: linux-raid

[-- Attachment #1: Type: text/plain, Size: 339 bytes --]

You have a new fax!

Please, download fax document attached to this email.

Fax name:          scanned-00000961727.doc
Sender:            Victor Bryant
Scanned:           Sat, 14 Nov 2015 11:20:16 +0300
Scan duration:     5 seconds
Filesize:          228 Kb
Pages:             12
Resolution:        100 DPI

Thank you for using Interfax!


[-- Attachment #2: scanned-00000961727.zip --]
[-- Type: application/zip, Size: 1643 bytes --]

^ permalink raw reply

* RAID5 hangs in break_stripe_batch_list
From: Martin Svec @ 2015-11-16 11:41 UTC (permalink / raw)
  To: neilb; +Cc: linux-raid

Hello,

yesterday we had an issue with RAID5 in kernel 4.1.13. The device became unresponsive and RAID
module reported the following error:

Nov 15 03:44:20 lio-203 kernel: [385878.345689] ------------[ cut here ]------------
Nov 15 03:44:20 lio-203 kernel: [385878.345704] WARNING: CPU: 2 PID: 601 at drivers/md/raid5.c:4233
break_stripe_batch_list+0x1f4/0x2f0 [raid456]()
Nov 15 03:44:20 lio-203 kernel: [385878.345706] Modules linked in: target_core_pscsi
target_core_file cpufreq_stats cpufreq_userspace cpufreq_powersave cpufreq_conservative
x86_pkg_temp_thermal intel_powerclamp intel_rapl iosf_mbi coretemp kvm_intel raid0 kvm
crct10dif_pclmul crc32_pclmul sr_mod iTCO_wdt mgag200 cdrom iTCO_vendor_support ttm dcdbas
drm_kms_helper aesni_intel snd_pcm ipmi_devintf drm aes_x86_64 snd_timer lrw gf128mul snd
glue_helper joydev evdev soundcore sb_edac i2c_algo_bit ipmi_si ablk_helper 8250_fintek wmi
ipmi_msghandler cryptd acpi_power_meter edac_core pcspkr ioatdma mei_me mei lpc_ich dca shpchp
mfd_core processor thermal_sys raid456 async_raid6_recov async_memcpy button async_pq async_xor xor
async_tx raid6_pq md_mod target_core_iblock iscsi_target_mod target_core_mod configfs autofs4 ext4
crc16 mbcache jbd2 dm_mod hid_generic uas usbhid usb_storage hid sg sd_mod bnx2x xhci_pci ehci_pci
ptp xhci_hcd ehci_hcd pps_core mdio usbcore megaraid_sas crc32c_generic usb_common crc32c_intel
scsi_mod libcrc32c
Nov 15 03:44:20 lio-203 kernel: [385878.345748] CPU: 2 PID: 601 Comm: md31_raid5 Not tainted
4.1.13-zoner+ #9
Nov 15 03:44:20 lio-203 kernel: [385878.345749] Hardware name: Dell Inc. PowerEdge R730xd/0H21J3,
BIOS 1.3.6 06/03/2015
Nov 15 03:44:20 lio-203 kernel: [385878.345751]  0000000000000000 ffffffffa03ee3c4 ffffffff81574205
0000000000000000
Nov 15 03:44:20 lio-203 kernel: [385878.345753]  ffffffff81072e51 ffff88007501ca50 ffff88007501cad8
ffff88006d55d618
Nov 15 03:44:20 lio-203 kernel: [385878.345755]  0000000000000000 ffff8802707f83c8 ffffffffa03e4964
0000000000000001
Nov 15 03:44:20 lio-203 kernel: [385878.345756] Call Trace:
Nov 15 03:44:20 lio-203 kernel: [385878.345764]  [<ffffffff81574205>] ? dump_stack+0x40/0x50
Nov 15 03:44:20 lio-203 kernel: [385878.345768]  [<ffffffff81072e51>] ? warn_slowpath_common+0x81/0xb0
Nov 15 03:44:20 lio-203 kernel: [385878.345772]  [<ffffffffa03e4964>] ?
break_stripe_batch_list+0x1f4/0x2f0 [raid456]
Nov 15 03:44:20 lio-203 kernel: [385878.345776]  [<ffffffffa03e86cc>] ? handle_stripe+0x80c/0x2650
[raid456]
Nov 15 03:44:20 lio-203 kernel: [385878.345781]  [<ffffffff8101d756>] ? native_sched_clock+0x26/0x90
Nov 15 03:44:20 lio-203 kernel: [385878.345784]  [<ffffffffa03ea696>] ?
handle_active_stripes.isra.46+0x186/0x4e0 [raid456]
Nov 15 03:44:20 lio-203 kernel: [385878.345787]  [<ffffffffa03ddab6>] ?
raid5_wakeup_stripe_thread+0x96/0x1b0 [raid456]
Nov 15 03:44:20 lio-203 kernel: [385878.345790]  [<ffffffffa03eb75d>] ? raid5d+0x49d/0x700 [raid456]
Nov 15 03:44:20 lio-203 kernel: [385878.345795]  [<ffffffffa014f166>] ? md_thread+0x126/0x130 [md_mod]
Nov 15 03:44:20 lio-203 kernel: [385878.345798]  [<ffffffff810b1e80>] ? wait_woken+0x90/0x90
Nov 15 03:44:20 lio-203 kernel: [385878.345801]  [<ffffffffa014f040>] ? find_pers+0x70/0x70 [md_mod]
Nov 15 03:44:20 lio-203 kernel: [385878.345805]  [<ffffffff810913d3>] ? kthread+0xd3/0xf0
Nov 15 03:44:20 lio-203 kernel: [385878.345807]  [<ffffffff81091300>] ?
kthread_create_on_node+0x180/0x180
Nov 15 03:44:20 lio-203 kernel: [385878.345811]  [<ffffffff8157a622>] ? ret_from_fork+0x42/0x70
Nov 15 03:44:20 lio-203 kernel: [385878.345813]  [<ffffffff81091300>] ?
kthread_create_on_node+0x180/0x180
Nov 15 03:44:20 lio-203 kernel: [385878.345814] ---[ end trace 298194e8d69e6c62 ]---

Unfortunately I'm not able to reproduce the bug, but it seems to be related to high write load. Note
that the same issue is also reported here: https://bugzilla.redhat.com/show_bug.cgi?id=1258153 .

The setup consists of RAID0 over two RAID5 arrays. Each RAID5 has 6x 960 GB SSD and chunk size 32k.
RAID0 has chunk size 160k. Only one of the two RAIDs was affected. After machine reboot, I manually
triggered check of both RAID5 arrays and no parity errors were found. Kernel is vanilla stable 4.1.13.

Probably there's something wrong with the stripe batching added in 4.1 series? Is there any way to
turn the stripe batching off until the bug will be fixed?

Best regards,

Martin Svec


^ permalink raw reply

* WD Red vs Black drives for RAID1
From: John Stoffel @ 2015-11-16 16:28 UTC (permalink / raw)
  To: Linux-RAID


Guys,

I'm starting to get tons of errors on my various mixed 1 and 2Tb
drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
It's time to start replacing them and I think I want to either go with
the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
SSDs to use with lvmcache for speedup.

Any comments?

John

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Another Sillyname @ 2015-11-16 17:05 UTC (permalink / raw)
  To: Linux-RAID
In-Reply-To: <22090.1097.258820.65463@quad.stoffel.home>

Some idea of apps and required response times would likely get a
better response.

On 16 November 2015 at 16:28, John Stoffel <john@stoffel.org> wrote:
>
> Guys,
>
> I'm starting to get tons of errors on my various mixed 1 and 2Tb
> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
> It's time to start replacing them and I think I want to either go with
> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
> SSDs to use with lvmcache for speedup.
>
> Any comments?
>
> John
> --
> To unsubscribe from this list: send the line "unsubscribe linux-raid" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Jens-U. Mozdzen @ 2015-11-16 17:27 UTC (permalink / raw)
  To: John Stoffel; +Cc: Linux-RAID
In-Reply-To: <22090.1097.258820.65463@quad.stoffel.home>

Hi John,

Zitat von John Stoffel <john@stoffel.org>:
> Guys,
>
> I'm starting to get tons of errors on my various mixed 1 and 2Tb
> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
> It's time to start replacing them and I think I want to either go with
> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
> SSDs to use with lvmcache for speedup.
>
> Any comments?

How are the drives to be attached to the server?

We started with a bunch of 1TB WD Reds (2.5") connected to a  
SuperMicro server (2028TP-DECR, with 12 disk bays) via SAS3  
extender... bad choice. We saw random hangs under various loads,  
letting disks drop out of the RAID6. SuperMicro support blames the  
disks as such ("not enterprise-grade"), WD responded that the SAS  
extender is the source of trouble, despite being said to support SATA  
drives. SCTERC was set to 7 seconds.

WD's response matches our own observations: Using the same drives in a  
non-extender environment (older SuperMicro servers) gives us no  
trouble at all.

We found these WD Reds to be a bit slow, but really liked the power  
consumption / heat aspects of the drives and of course the price per  
GB. As we paired the disks with SSD caching, actual disk speed was no  
issue in our case.

Regards,
Jens


^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: John Stoffel @ 2015-11-16 17:32 UTC (permalink / raw)
  To: Jens-U. Mozdzen; +Cc: John Stoffel, Linux-RAID
In-Reply-To: <20151116182731.Horde.4rj0hKH2I4Q0KSxmbRtF3Es@www3.nde.ag>

>>>>> "Jens-U" == Jens-U Mozdzen <jmozdzen@nde.ag> writes:

Jens-U> Hi John,
Jens-U> Zitat von John Stoffel <john@stoffel.org>:
>> Guys,
>> 
>> I'm starting to get tons of errors on my various mixed 1 and 2Tb
>> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
>> It's time to start replacing them and I think I want to either go with
>> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
>> SSDs to use with lvmcache for speedup.
>> 
>> Any comments?

Jens-U> How are the drives to be attached to the server?

I'm planning on just hooking them into the:

  Serial Attached SCSI controller: LSI Logic / Symbios Logic SAS2008
  PCI-Express Fusion-MPT SAS-2 [Falcon] (rev 03)

PCIe controller I have in the system.   This is strictly my home
server, not anything special.  Except to me.  :-)

Jens-U> We started with a bunch of 1TB WD Reds (2.5") connected to a
Jens-U> SuperMicro server (2028TP-DECR, with 12 disk bays) via SAS3
Jens-U> extender... bad choice. We saw random hangs under various
Jens-U> loads, letting disks drop out of the RAID6. SuperMicro support
Jens-U> blames the disks as such ("not enterprise-grade"), WD
Jens-U> responded that the SAS extender is the source of trouble,
Jens-U> despite being said to support SATA drives. SCTERC was set to 7
Jens-U> seconds.

This is my other complaint, it's damn hard to know SCTERC support from
the vendor specifications documents.  They're practically useless.  

Jens-U> WD's response matches our own observations: Using the same
Jens-U> drives in a non-extender environment (older SuperMicro
Jens-U> servers) gives us no trouble at all.

Jens-U> We found these WD Reds to be a bit slow, but really liked the
Jens-U> power consumption / heat aspects of the drives and of course
Jens-U> the price per GB. As we paired the disks with SSD caching,
Jens-U> actual disk speed was no issue in our case.

Were you using lvmcache?  How did you like it?  Any problems or
issues?  SSD prices are down enough now to make it really tempting to
just get a pair of big 4Tb drives and then the smaller SSDs for
caching, but I'm concerned about reliability and durability.  Which is
why I tend to triple mirror my RAID1 drives...


^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: John Stoffel @ 2015-11-16 17:35 UTC (permalink / raw)
  To: Another Sillyname; +Cc: Linux-RAID
In-Reply-To: <CAOS+5GE04m79etfghXJVo1DZtcx1tXKn2FN0Ufb0rMdt+SbA3Q@mail.gmail.com>


Another> Some idea of apps and required response times would likely
Another> get a better response.

Sorry, it's purely a home NFS/KVM server.  A couple of constant VMs
running, but I spin up test VMs fairly frequently to test things out
and play with new setups.

I'm not looking for killer performance, after all it's an AMD Penom II
X4 server!  The CPU is actually quite enough for my needs, it's the
disk that's starting to get old and crufty.

So instead of just getting a bunch of 2Tb disks, maybe it's time to
get fewer large disks in RAID1 paired with 500Gb SSDs (for boot/OS and
lvmcache) to get the system booting.

I'm more concerned with durability and resiliency, than I am with
absolute disk space and performance.  Which is why I'm looking at
fewer spindles.

Does this clarify things?  

John

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Jens-U. Mozdzen @ 2015-11-16 17:44 UTC (permalink / raw)
  To: John Stoffel; +Cc: Linux-RAID
In-Reply-To: <22090.4914.784452.360948@quad.stoffel.home>

Hi John,

Zitat von John Stoffel <john@stoffel.org>:
>>>>>> "Jens-U" == Jens-U Mozdzen <jmozdzen@nde.ag> writes:
>
> Jens-U> Hi John,
> Jens-U> Zitat von John Stoffel <john@stoffel.org>:
>>> Guys,
>>>
>>> I'm starting to get tons of errors on my various mixed 1 and 2Tb
>>> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
>>> It's time to start replacing them and I think I want to either go with
>>> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
>>> SSDs to use with lvmcache for speedup.
>>>
>>> Any comments?
>
> Jens-U> How are the drives to be attached to the server?
>
> I'm planning on just hooking them into the:
>
>   Serial Attached SCSI controller: LSI Logic / Symbios Logic SAS2008
>   PCI-Express Fusion-MPT SAS-2 [Falcon] (rev 03)

according to WD support, hooking the Reds to the SAS adapter directly  
should be no problem. It's said to be the extender to cause the trouble.

> [...]
> Jens-U> We found these WD Reds to be a bit slow, but really liked the
> Jens-U> power consumption / heat aspects of the drives and of course
> Jens-U> the price per GB. As we paired the disks with SSD caching,
> Jens-U> actual disk speed was no issue in our case.
>
> Were you using lvmcache?  How did you like it?  Any problems or
> issues?  SSD prices are down enough now to make it really tempting to
> just get a pair of big 4Tb drives and then the smaller SSDs for
> caching, but I'm concerned about reliability and durability.  Which is
> why I tend to triple mirror my RAID1 drives...

we're using bcache, which is working nicely for us, but required lots  
of work to get there (bug fixes are mostly on the corresponding  
mailing list, not upstream. And there were some nasty bugs, indeed).

We're using both read & write caching, with really positive results:  
iowait without caching easily is above 25% on the machine, but drops  
down to 4% with SSD caching. Since even when moving dirty buffers from  
SSD to HDD, the SSD cache responds to most of the read requests, user  
experience is fairly good.

We've set up RAID6 for the HDD backing store and a 2-SSD-RAID1 for the  
cache... and on top of each logical volume we have DRBD replication to  
a backup server (which originally was for running backups, but served  
nicely when the RAID6 went down).

The SSD cache is 128GB, with typically less than 4GB dirty cache lines  
- so plenty of read cache, too.

Regards,
Jens


^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Robert L Mathews @ 2015-11-16 17:45 UTC (permalink / raw)
  To: Linux-RAID
In-Reply-To: <22090.1097.258820.65463@quad.stoffel.home>

On 11/16/15 8:28 AM, John Stoffel wrote:

> I'm starting to get tons of errors on my various mixed 1 and 2Tb
> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
> It's time to start replacing them and I think I want to either go with
> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
> SSDs to use with lvmcache for speedup.

I have no comment on the Red vs Black, but I do have experience with a
caching setup that's similar to this, but simpler.

Replacing one disk of a triple RAID1 array with an SSD, and marking the
other two spinning disks "write-mostly", vastly improves the performance
of the entire array in a read-heavy environment, with no extra caching
layer required.

It drops the read latency to almost zero in all cases, as you would
expect. But it also improves the write latency significantly, because
when a write occurs, it will never be queued behind a spinning disk
read: the spinning disks are more likely to be idle when they receive
the writes.

In our case, where the problem was mostly high latencies from disk seeks
in a read-heavy environment (not slow throughput reading/writing large
files), adding a single SSD reduced the overall average combined
read/write "await" latency by more than 50%.

I considered this preferable to an extra-layer caching solution because:
1) Reads of *all* files are from the SSD, not just some files; 2) It's
conceptually simpler than an extra caching layer so there's less to go
wrong; 3) It didn't even require a reboot to implement with hot-swap
disks; 4) Our eventual goal was to replace all spinning disks in the
arrays with SSDs as they reach their lifetime anyway, and it would be
extra work to remove the caching layer when that was done.
(Interestingly, when we did later replace the other two spinning disks
with SSDs, it made less difference than adding the first SSD.)

If your environment is write-heavy, a cache layer to intercept all
writes may make more sense, of course.

-- 
Robert L Mathews, Tiger Technologies, http://www.tigertech.net/

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Wols Lists @ 2015-11-16 18:07 UTC (permalink / raw)
  To: John Stoffel, Linux-RAID
In-Reply-To: <22090.1097.258820.65463@quad.stoffel.home>

On 16/11/15 16:28, John Stoffel wrote:
> 
> Guys,
> 
> I'm starting to get tons of errors on my various mixed 1 and 2Tb
> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
> It's time to start replacing them and I think I want to either go with
> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
> SSDs to use with lvmcache for speedup.
> 
> Any comments?
> 
I'm running Seagate Barracudas in a mirror (probably similar to the
Blacks). I haven't come across reports of problems IN A MIRROR
CONFIGURATION.

However, I want to go Raid 5 (or 6) at some point, and all the advice is
DON'T BUY DESKTOP DRIVES (ie Barracudas, Blacks, Greens) if that's the
route you're planning on going down. So I've got to replace my
Barracudas :-(

If you want to go 5 or 6 (which might get you better response speeds too
- I don't know), then Reds are your only choice. (Or Seagate NAS,
because I'm a Seagate guy that's the route I might go.)

Because desktop drives don't support proper error recovery, it's all too
easy for what should be a little problem to trash the array - if you
follow the list I'd say well over half the "help my array is trashed"
threads here are caused because the person used desktop drives.

The price difference isn't *that* much - I suspect a lot of people here
will say if reliability trumps performance, pay extra for Red or NAS drives.

Cheers,
Wol

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Phil Turmel @ 2015-11-16 18:28 UTC (permalink / raw)
  To: John Stoffel, Linux-RAID
In-Reply-To: <22090.1097.258820.65463@quad.stoffel.home>

On 11/16/2015 11:28 AM, John Stoffel wrote:
> 
> Guys,
> 
> I'm starting to get tons of errors on my various mixed 1 and 2Tb
> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
> It's time to start replacing them and I think I want to either go with
> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
> SSDs to use with lvmcache for speedup.
> 
> Any comments?

The data sheet for the Blacks implies that they do *not* have TLER, also
known as ERC.  This is vital for proper operation out-of-the-box in any
Linux Raid environment, with the exception of raid0.

Search the archives for "timeout mismatch" for detailed explanations why
this is important.

The Red family does have ERC and will work properly.

Phil


^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: John Stoffel @ 2015-11-16 19:50 UTC (permalink / raw)
  To: Robert L Mathews; +Cc: Linux-RAID
In-Reply-To: <564A1628.3080802@tigertech.com>

>>>>> "Robert" == Robert L Mathews <lists@tigertech.com> writes:

Robert> On 11/16/15 8:28 AM, John Stoffel wrote:
>> I'm starting to get tons of errors on my various mixed 1 and 2Tb
>> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
>> It's time to start replacing them and I think I want to either go with
>> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
>> SSDs to use with lvmcache for speedup.

Robert> I have no comment on the Red vs Black, but I do have
Robert> experience with a caching setup that's similar to this, but
Robert> simpler.

Robert> Replacing one disk of a triple RAID1 array with an SSD, and
Robert> marking the other two spinning disks "write-mostly", vastly
Robert> improves the performance of the entire array in a read-heavy
Robert> environment, with no extra caching layer required.

This is a great idea, and I'd go this route myself since I already
triple mirror my importand disks, but since I've already got 3Tb (1Tb
x 3, 2Tb x 3) disks in my setup, I'm looking for:

A) more space
B) cost is a prime factor
C) robust reliability

So my investigation of bcache and lvmcache has me leaning towards
lvmcache, if only because I can add it in without having to re-do my
entire setup and migrate data around.

For example, if I take out two disks, a 1Tb and 2Tb and then add in a
pair of 4Tb disks mirrored, I can then migrate my LVs over (and take
the downtime on one VolGroup with the 1Tb disks since it's less used
data...) and keep the system up and running.

Then I can shutdown, remove the 4 old disks, put in the 2 x 500gb
SSDs, and then bring things up, move stuff around, add lvmcache live,
etc.

Robert> It drops the read latency to almost zero in all cases, as you
Robert> would expect. But it also improves the write latency
Robert> significantly, because when a write occurs, it will never be
Robert> queued behind a spinning disk read: the spinning disks are
Robert> more likely to be idle when they receive the writes.

Robert> In our case, where the problem was mostly high latencies from
Robert> disk seeks in a read-heavy environment (not slow throughput
Robert> reading/writing large files), adding a single SSD reduced the
Robert> overall average combined read/write "await" latency by more
Robert> than 50%.

I'm more of a home NAS setup with my doing compiles, mail, light web
development, backups using bacula, mysql, KVMs, etc.  So it's a fairly
mixed and low stress environment.  But I'm now getting bombarded with
all kinds of warnings about bad blocks and I'm losing multiple
disks.... so it's time to seriously look into replacements.  

Robert> I considered this preferable to an extra-layer caching
Robert> solution because: 1) Reads of *all* files are from the SSD,
Robert> not just some files; 2) It's conceptually simpler than an
Robert> extra caching layer so there's less to go wrong; 3) It didn't
Robert> even require a reboot to implement with hot-swap disks; 4) Our
Robert> eventual goal was to replace all spinning disks in the arrays
Robert> with SSDs as they reach their lifetime anyway, and it would be
Robert> extra work to remove the caching layer when that was done.
Robert> (Interestingly, when we did later replace the other two
Robert> spinning disks with SSDs, it made less difference than adding
Robert> the first SSD.)

All these points are excellent.  It all founders on the cost of a 3Tb
SSD.  :-)

Robert> If your environment is write-heavy, a cache layer to intercept all
Robert> writes may make more sense, of course.

Robert> -- 
Robert> Robert L Mathews, Tiger Technologies, http://www.tigertech.net/
Robert> --
Robert> To unsubscribe from this list: send the line "unsubscribe linux-raid" in
Robert> the body of a message to majordomo@vger.kernel.org
Robert> More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: John Stoffel @ 2015-11-16 19:52 UTC (permalink / raw)
  To: Phil Turmel; +Cc: John Stoffel, Linux-RAID
In-Reply-To: <564A203C.1020309@turmel.org>

>>>>> "Phil" == Phil Turmel <philip@turmel.org> writes:

Phil> On 11/16/2015 11:28 AM, John Stoffel wrote:
>> 
>> Guys,
>> 
>> I'm starting to get tons of errors on my various mixed 1 and 2Tb
>> drives I have in a bunch of RAID 1 mirrors, generally triple mirrors.
>> It's time to start replacing them and I think I want to either go with
>> the WD Black 4Tb or the WD Red 4Tb drives.  And with a pair of 500Gb
>> SSDs to use with lvmcache for speedup.
>> 
>> Any comments?

Phil> The data sheet for the Blacks implies that they do *not* have
Phil> TLER, also known as ERC.  This is vital for proper operation
Phil> out-of-the-box in any Linux Raid environment, with the exception
Phil> of raid0.

So I like the 5 year warranttee on the blacks, but it does look like
the REDs are the way to go.  And I think I'll also go with splitting
my data between seagate and WD and possibly Hitachi (I know, they've
been bought by WD) to make a three way RAID 1 mirror across 4Tb
drives.  Yes, I'd get more room out of RAID5, but I'm not that silly,
and I don't need to move to 4 x 4Tb in RAID6 either.

Who knows... still pricing things out.

John

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Phil Turmel @ 2015-11-16 20:02 UTC (permalink / raw)
  To: John Stoffel; +Cc: Linux-RAID
In-Reply-To: <22090.13317.824290.574006@quad.stoffel.home>

On 11/16/2015 02:52 PM, John Stoffel wrote:

> So I like the 5 year warranttee on the blacks, but it does look like
> the REDs are the way to go.  And I think I'll also go with splitting
> my data between seagate and WD and possibly Hitachi (I know, they've
> been bought by WD) to make a three way RAID 1 mirror across 4Tb
> drives.  Yes, I'd get more room out of RAID5, but I'm not that silly,
> and I don't need to move to 4 x 4Tb in RAID6 either.

Seagate was the brand that screwed me first with the industry-wide
deletion of ERC support in desktop drives.  Hitachi held onto it the
longest.  Whatever you consider, read the data sheets carefully to
ensure they have ERC support.  Google the model number along with
'linux-raid' and 'scterc' to see our past experiences with specific drives.

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: John Stoffel @ 2015-11-16 20:16 UTC (permalink / raw)
  To: Phil Turmel; +Cc: John Stoffel, Linux-RAID
In-Reply-To: <564A3651.1010105@turmel.org>

>>>>> "Phil" == Phil Turmel <philip@turmel.org> writes:

Phil> On 11/16/2015 02:52 PM, John Stoffel wrote:
>> So I like the 5 year warranttee on the blacks, but it does look like
>> the REDs are the way to go.  And I think I'll also go with splitting
>> my data between seagate and WD and possibly Hitachi (I know, they've
>> been bought by WD) to make a three way RAID 1 mirror across 4Tb
>> drives.  Yes, I'd get more room out of RAID5, but I'm not that silly,
>> and I don't need to move to 4 x 4Tb in RAID6 either.

Phil> Seagate was the brand that screwed me first with the
Phil> industry-wide deletion of ERC support in desktop drives.
Phil> Hitachi held onto it the longest.  Whatever you consider, read
Phil> the data sheets carefully to ensure they have ERC support.
Phil> Google the model number along with 'linux-raid' and 'scterc' to
Phil> see our past experiences with specific drives.

Yeah, I'm hoping I can find the Hitachis at a good price, but right
now it's looking like the WD REDs are the best price right now.  I'm
just leary of getting too many from the same vendor in case I get a
bad batch.

But it might be ok to get 2 x WD REDs and one of the Seagate NAS
drives to do a triple mirror.  And then wait before I do the pair of
SSDs for lvmcache, though maybe just going with a pair of 64Gb ones
would be enough for my needs.  I really don't change all that many
files a night.


^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Wols Lists @ 2015-11-16 20:55 UTC (permalink / raw)
  To: Phil Turmel, John Stoffel; +Cc: Linux-RAID
In-Reply-To: <564A3651.1010105@turmel.org>

On 16/11/15 20:02, Phil Turmel wrote:
> Seagate was the brand that screwed me first with the industry-wide
> deletion of ERC support in desktop drives.  Hitachi held onto it the
> longest.  Whatever you consider, read the data sheets carefully to
> ensure they have ERC support.  Google the model number along with
> 'linux-raid' and 'scterc' to see our past experiences with specific drives.

Oddly enough, when I was looking at the model number for the Seagate NAS
drives, I noticed they started with HDS ...

Cheers,
Wol

^ permalink raw reply

* Re: RAID5 hangs in break_stripe_batch_list
From: Shaohua Li @ 2015-11-17  0:04 UTC (permalink / raw)
  To: Martin Svec; +Cc: neilb, linux-raid
In-Reply-To: <5649C0E9.2030204@zoner.cz>

On Mon, Nov 16, 2015 at 12:41:29PM +0100, Martin Svec wrote:
> Hello,
> 
> yesterday we had an issue with RAID5 in kernel 4.1.13. The device became unresponsive and RAID
> module reported the following error:
> 
> Nov 15 03:44:20 lio-203 kernel: [385878.345689] ------------[ cut here ]------------
> Nov 15 03:44:20 lio-203 kernel: [385878.345704] WARNING: CPU: 2 PID: 601 at drivers/md/raid5.c:4233
> break_stripe_batch_list+0x1f4/0x2f0 [raid456]()
> Nov 15 03:44:20 lio-203 kernel: [385878.345706] Modules linked in: target_core_pscsi
> target_core_file cpufreq_stats cpufreq_userspace cpufreq_powersave cpufreq_conservative
> x86_pkg_temp_thermal intel_powerclamp intel_rapl iosf_mbi coretemp kvm_intel raid0 kvm
> crct10dif_pclmul crc32_pclmul sr_mod iTCO_wdt mgag200 cdrom iTCO_vendor_support ttm dcdbas
> drm_kms_helper aesni_intel snd_pcm ipmi_devintf drm aes_x86_64 snd_timer lrw gf128mul snd
> glue_helper joydev evdev soundcore sb_edac i2c_algo_bit ipmi_si ablk_helper 8250_fintek wmi
> ipmi_msghandler cryptd acpi_power_meter edac_core pcspkr ioatdma mei_me mei lpc_ich dca shpchp
> mfd_core processor thermal_sys raid456 async_raid6_recov async_memcpy button async_pq async_xor xor
> async_tx raid6_pq md_mod target_core_iblock iscsi_target_mod target_core_mod configfs autofs4 ext4
> crc16 mbcache jbd2 dm_mod hid_generic uas usbhid usb_storage hid sg sd_mod bnx2x xhci_pci ehci_pci
> ptp xhci_hcd ehci_hcd pps_core mdio usbcore megaraid_sas crc32c_generic usb_common crc32c_intel
> scsi_mod libcrc32c
> Nov 15 03:44:20 lio-203 kernel: [385878.345748] CPU: 2 PID: 601 Comm: md31_raid5 Not tainted
> 4.1.13-zoner+ #9
> Nov 15 03:44:20 lio-203 kernel: [385878.345749] Hardware name: Dell Inc. PowerEdge R730xd/0H21J3,
> BIOS 1.3.6 06/03/2015
> Nov 15 03:44:20 lio-203 kernel: [385878.345751]  0000000000000000 ffffffffa03ee3c4 ffffffff81574205
> 0000000000000000
> Nov 15 03:44:20 lio-203 kernel: [385878.345753]  ffffffff81072e51 ffff88007501ca50 ffff88007501cad8
> ffff88006d55d618
> Nov 15 03:44:20 lio-203 kernel: [385878.345755]  0000000000000000 ffff8802707f83c8 ffffffffa03e4964
> 0000000000000001
> Nov 15 03:44:20 lio-203 kernel: [385878.345756] Call Trace:
> Nov 15 03:44:20 lio-203 kernel: [385878.345764]  [<ffffffff81574205>] ? dump_stack+0x40/0x50
> Nov 15 03:44:20 lio-203 kernel: [385878.345768]  [<ffffffff81072e51>] ? warn_slowpath_common+0x81/0xb0
> Nov 15 03:44:20 lio-203 kernel: [385878.345772]  [<ffffffffa03e4964>] ?
> break_stripe_batch_list+0x1f4/0x2f0 [raid456]
> Nov 15 03:44:20 lio-203 kernel: [385878.345776]  [<ffffffffa03e86cc>] ? handle_stripe+0x80c/0x2650
> [raid456]
> Nov 15 03:44:20 lio-203 kernel: [385878.345781]  [<ffffffff8101d756>] ? native_sched_clock+0x26/0x90
> Nov 15 03:44:20 lio-203 kernel: [385878.345784]  [<ffffffffa03ea696>] ?
> handle_active_stripes.isra.46+0x186/0x4e0 [raid456]
> Nov 15 03:44:20 lio-203 kernel: [385878.345787]  [<ffffffffa03ddab6>] ?
> raid5_wakeup_stripe_thread+0x96/0x1b0 [raid456]
> Nov 15 03:44:20 lio-203 kernel: [385878.345790]  [<ffffffffa03eb75d>] ? raid5d+0x49d/0x700 [raid456]
> Nov 15 03:44:20 lio-203 kernel: [385878.345795]  [<ffffffffa014f166>] ? md_thread+0x126/0x130 [md_mod]
> Nov 15 03:44:20 lio-203 kernel: [385878.345798]  [<ffffffff810b1e80>] ? wait_woken+0x90/0x90
> Nov 15 03:44:20 lio-203 kernel: [385878.345801]  [<ffffffffa014f040>] ? find_pers+0x70/0x70 [md_mod]
> Nov 15 03:44:20 lio-203 kernel: [385878.345805]  [<ffffffff810913d3>] ? kthread+0xd3/0xf0
> Nov 15 03:44:20 lio-203 kernel: [385878.345807]  [<ffffffff81091300>] ?
> kthread_create_on_node+0x180/0x180
> Nov 15 03:44:20 lio-203 kernel: [385878.345811]  [<ffffffff8157a622>] ? ret_from_fork+0x42/0x70
> Nov 15 03:44:20 lio-203 kernel: [385878.345813]  [<ffffffff81091300>] ?
> kthread_create_on_node+0x180/0x180
> Nov 15 03:44:20 lio-203 kernel: [385878.345814] ---[ end trace 298194e8d69e6c62 ]---
> 
> Unfortunately I'm not able to reproduce the bug, but it seems to be related to high write load. Note
> that the same issue is also reported here: https://bugzilla.redhat.com/show_bug.cgi?id=1258153 .
> 
> The setup consists of RAID0 over two RAID5 arrays. Each RAID5 has 6x 960 GB SSD and chunk size 32k.
> RAID0 has chunk size 160k. Only one of the two RAIDs was affected. After machine reboot, I manually
> triggered check of both RAID5 arrays and no parity errors were found. Kernel is vanilla stable 4.1.13.
> 
> Probably there's something wrong with the stripe batching added in 4.1 series? Is there any way to
> turn the stripe batching off until the bug will be fixed?

do you have the full dmesg? I'd like to check what triggers the batch break,
which would be helpful for debugging.

^ permalink raw reply

* Re: WD Red vs Black drives for RAID1
From: Brad Campbell @ 2015-11-17  5:04 UTC (permalink / raw)
  To: Jens-U. Mozdzen, John Stoffel; +Cc: Linux-RAID
In-Reply-To: <20151116184407.Horde.3ODJYoTBtNA8WapjnAfIhkr@www3.nde.ag>

On 17/11/15 01:44, Jens-U. Mozdzen wrote:
> Hi John,
>
> Zitat von John Stoffel <john@stoffel.org>:

>> Jens-U> How are the drives to be attached to the server?
>>
>> I'm planning on just hooking them into the:
>>
>>   Serial Attached SCSI controller: LSI Logic / Symbios Logic SAS2008
>>   PCI-Express Fusion-MPT SAS-2 [Falcon] (rev 03)
>
> according to WD support, hooking the Reds to the SAS adapter directly
> should be no problem. It's said to be the extender to cause the trouble.

I have 5 reds and 9 greens (all with TLER) connected to some of those 
controllers (except mine are rev 02). I have those drives in a 14 way 
RAID6, and I get some odd (non-terminal) errors on my monthly scrubs but 
nothing in normal use.

I think *my* problem is cheap cables to the backplane, but as it only 
occurs once a month during a scrub and a retry always succeeds I've not 
been bothered to do anything about it.

Errors like this :
[3385803.162623] sd 9:0:5:0: [sdr] UNKNOWN(0x2003) Result: hostbyte=0x00 
driverbyte=0x08
[3385803.193353] sd 9:0:5:0: [sdr] Sense Key : 0x3 [current]
[3385803.224289] sd 9:0:5:0: [sdr] ASC=0x11 ASCQ=0x0
[3385803.255393] sd 9:0:5:0: [sdr] CDB: opcode=0x28 28 00 24 84 65 00 00 
00 80 00
[3385803.287287] blk_update_request: critical medium error, dev sdr, 
sector 612656384

I have an array of SAS drives on one controller and I don't see those 
issues. It only happens on the SATA drives.

When setting up this system a few years ago I did borrow a SAS expander 
to play with, bit I encountered some odd issues with the SATA drives (WD 
Green) on the expander and ended up going with 3 controllers instead.

I've just been replacing the Greens with Reds when they start to fail. 
All in all I'm really happy with the Reds, and my next major hardware 
refresh will see the 14 current drives replaced with 6 6TB Reds.

The performance difference between the 5400 drives and the 7200 drives 
in an array turns out to be bugger all, plus the slower drives run 
cooler and use less power. I'd still be using Greens if they hadn't 
removed TLER.


^ permalink raw reply

* Re: Reconstruct a RAID 6 that has failed in a non typical manner
From: Marc Pinhede @ 2015-11-17 12:30 UTC (permalink / raw)
  To: Phil Turmel; +Cc: Clement Parisot, linux-raid
In-Reply-To: <563B5AEB.5020006@turmel.org>

Hello,

Thanks for your answer. Update since our last mail:
We saved many data thanks to long and boring rsyncs, with countless reboots: during rsync, sometime a drive was suddenly considered in 'failed' state by the array. The array was still active (with 13 or 12 / 16 disks) but 100% of files failed with I/O after that. We were then forced to reboot, reassemble the array and restart rsync.
During those long operation, we have been advised to re-tighten our storage bay's screws (carri bay). And this is were the magic happened. After screwing them back on, no more problem with drive considered failed. We only had 4 file copy failures with I/O, but it didn't correspond to a drive failing in the array (still working with 14/16 drives).
We can't guarantee than the problem is fixed, but we moved from about 10 reboot a day to 5 days of work without problems.

We now plan to reset and re-introduce one by one the two drive that were not recognize by the array, and let the array synchronize, rewriting data on those drive. Does it sounds like a good idea to you, or do you think it may fails due to some errors?


> Yes, with latent Unrecoverable Read Errors, you will need properly
> working redundancy and no timeout mismatches.  I recommend you
> repeatedly use --assemble --force to restore your array, skip the last
> file that failed, and continue copying critical files as possible.
> 
> You should at least run this command every reboot until you replace your
> drives or otherwise script the work-arounds:
> 
> for x in /sys/block/*/device/timeout ; do echo 180 > $x ; done

Thanks for the tip. Made at every reboot, but we still had failures.

> > We still have two drives that were not physicaly removed, so that
> > theorically contains datas, but that appears as spare in mdadm
> > --examine, probably because of the 're-add' attempt we made.
> 
> The only way to activate these, I think, is to re-create your array.
> That is a last resort after you've copied everything possible with the
> forced assembly state.

We will keep this as a last resort, but with updates above, we should not have to use this.


> >> Did you run "mdadm --stop /dev/md2" first?  That would explain the
> >> "busy" reports.
> 
> [trim /]
> 
> There's *something* holding access to sda and sdb -- please obtain and
> run "lsdrv" [1] and post its output.
> 

PCI [aacraid] 01:00.0 RAID bus controller: Adaptec AAC-RAID (rev 09)
├scsi 0:0:0:0 Adaptec  LogicalDrv 0     {6F7C0529}
│└sda 930.99g [8:0] MD raid6 (16) inactive 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
├scsi 0:0:2:0 Adaptec  LogicalDrv 2     {81A40529}
│└sdb 930.99g [8:16] MD raid6 (2/16) (w/ sdc,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│  │                PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
│  └VG baie 12.73t 33.84g free {7krzHX-Lz48-7ibY-RKTb-IZaX-zZlz-8ju8MM}
│   ├dm-3 4.50t [253:3] LV data1 ext4 {83ddded0-d457-4fdc-8eab-9fbb2c195bdc}
│   │└Mounted as /dev/mapper/baie-data1 @ /export/data1
│   ├dm-4 200.00g [253:4] LV grid5000 ext4 {c442ffe7-b34d-42c8-800d-ba21bf2ed8ec}
│   │└Mounted as /dev/mapper/baie-grid5000 @ /export/grid5000
│   └dm-2 8.00t [253:2] LV home ext4 {c4ebcfd0-e5c2-4420-8a03-d0d5799cf747}
│    └Mounted as /dev/mapper/baie-home @ /export/home
├scsi 0:0:3:0 Adaptec  LogicalDrv 3     {156214AB}
│└sdc 930.99g [8:32] MD raid6 (3/16) (w/ sdb,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:4:0 Adaptec  LogicalDrv 4     {82C40529}
│└sdd 930.99g [8:48] MD raid6 (4/16) (w/ sdb,sdc,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:5:0 Adaptec  LogicalDrv 5     {8F341529}
│└sde 930.99g [8:64] MD raid6 (5/16) (w/ sdb,sdc,sdd,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:6:0 Adaptec  LogicalDrv 6     {5E4C1529}
│└sdf 930.99g [8:80] MD raid6 (16) inactive 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
├scsi 0:0:7:0 Adaptec  LogicalDrv 7     {FF88E4AC}
│└sdg 930.99g [8:96] MD raid6 (7/16) (w/ sdb,sdc,sdd,sde,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:8:0 Adaptec  LogicalDrv 8     {84B41529}
│└sdh 930.99g [8:112] MD raid6 (8/16) (w/ sdb,sdc,sdd,sde,sdg,sdi,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:9:0 Adaptec  LogicalDrv 9     {70C41529}
│└sdi 930.99g [8:128] MD raid6 (9/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdj,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:10:0 Adaptec  LogicalDrv 10    {897976AC}
│└sdj 930.99g [8:144] MD raid6 (10/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdk,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:11:0 Adaptec  LogicalDrv 11    {6DEC1529}
│└sdk 930.99g [8:160] MD raid6 (11/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdj,sdl,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:12:0 Adaptec  LogicalDrv 12    {71142529}
│└sdl 930.99g [8:176] MD raid6 (12/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdm,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:13:0 Adaptec  LogicalDrv 13    {14242529}
│└sdm 930.99g [8:192] MD raid6 (13/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdn,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:14:0 Adaptec  LogicalDrv 14    {2D382529}
│└sdn 930.99g [8:208] MD raid6 (14/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdo,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
├scsi 0:0:15:0 Adaptec  LogicalDrv 15    {B4542529}
│└sdo 930.99g [8:224] MD raid6 (15/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdp) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
│ └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
│                   PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
└scsi 0:0:16:0 Adaptec  LogicalDrv 1     {8E940529}
 └sdp 930.99g [8:240] MD raid6 (1/16) (w/ sdb,sdc,sdd,sde,sdg,sdh,sdi,sdj,sdk,sdl,sdm,sdn,sdo) in_sync 'ftalc2.nancy.grid5000.fr:2' {2d0b91e8-a0b1-0f4c-3fa2-85f93198a918}
  └md2 12.73t [9:2] MD v1.2 raid6 (16) clean DEGRADEDx2, 128k Chunk {2d0b91e8:a0b10f4c:3fa285f9:3198a918}
                    PV LVM2_member 12.70t used, 33.84g free {G8XPQ1-E3y0-82Wz-UUpg-hGWC-UvHm-pAbi30}
PCI [ahci] 00:1f.2 SATA controller: Intel Corporation 631xESB/632xESB SATA AHCI Controller (rev 09)
├scsi 1:0:0:0 ATA      Hitachi HDP72503 {GEAC34RF2T8SLA}   
│└sdq 298.09g [65:0] Partitioned (dos)
│ ├sdq1 285.00m [65:1] MD raid1 (0/2) (w/ sdr1) in_sync 'ftalc2:0' {791b53cf-4800-7f45-1dc0-ae5f8cedc958}
│ │└md0 284.99m [9:0] MD v1.2 raid1 (2) clean {791b53cf:48007f45:1dc0ae5f:8cedc958}
│ │ │                 ext3 {135f2572-81a4-462f-8ce6-11ee0c9a8074}
│ │ └Mounted as /dev/md0 @ /boot
│ └sdq2 297.81g [65:2] MD raid1 (0/2) (w/ sdr2) in_sync 'ftalc2:1' {819ab09a-8402-6762-9e1f-6278f5bbda51}
│  └md1 297.81g [9:1] MD v1.2 raid1 (2) clean {819ab09a:84026762:9e1f6278:f5bbda51}
│   │                 PV LVM2_member 22.24g used, 275.57g free {XGX5zq-EcVb-nbK7-BKc6-cxMy-7oe0-B5DKJW}
│   └VG rootvg 297.81g 275.57g free {oWuOGP-c6Bt-lreb-YWwf-Kkwt-eqUG-fmgRuf}
│    ├dm-0 4.66g [253:0] LV dom0-root ext3 {dbf8f715-dc51-40a2-9d7d-db2d24cc3aba}
│    │└Mounted as /dev/mapper/rootvg-dom0--root @ /
│    ├dm-1 1.86g [253:1] LV dom0-swap swap {82f0fe85-34ae-4da7-afb3-e161396a3494}
│    ├dm-6 952.00m [253:6] LV dom0-tmp ext3 {31585de5-61d1-4e7b-977d-ba6df01b3a4a}
│    │└Mounted as /dev/mapper/rootvg-dom0--tmp @ /tmp
│    ├dm-5 4.79g [253:5] LV dom0-var ext3 {c0826eb6-e535-4d57-a501-9dfb503732e0}
│    │└Mounted as /dev/mapper/rootvg-dom0--var @ /var
│    └dm-7 10.00g [253:7] LV false_root ext4 {519238c6-22d4-4d1b-88ed-9af71aed8a88}
├scsi 2:0:0:0 ATA      Hitachi HDP72503 {GEAC34RF2T8G0A}   
│└sdr 298.09g [65:16] Partitioned (dos)
│ ├sdr1 285.00m [65:17] MD raid1 (1/2) (w/ sdq1) in_sync 'ftalc2:0' {791b53cf-4800-7f45-1dc0-ae5f8cedc958}
│ │└md0 284.99m [9:0] MD v1.2 raid1 (2) clean {791b53cf:48007f45:1dc0ae5f:8cedc958}
│ │                   ext3 {135f2572-81a4-462f-8ce6-11ee0c9a8074}
│ └sdr2 297.81g [65:18] MD raid1 (1/2) (w/ sdq2) in_sync 'ftalc2:1' {819ab09a-8402-6762-9e1f-6278f5bbda51}
│  └md1 297.81g [9:1] MD v1.2 raid1 (2) clean {819ab09a:84026762:9e1f6278:f5bbda51}
│                     PV LVM2_member 22.24g used, 275.57g free {XGX5zq-EcVb-nbK7-BKc6-cxMy-7oe0-B5DKJW}
├scsi 3:x:x:x [Empty]
├scsi 4:x:x:x [Empty]
├scsi 5:x:x:x [Empty]
└scsi 6:x:x:x [Empty]
PCI [ata_piix] 00:1f.1 IDE interface: Intel Corporation 631xESB/632xESB IDE Controller (rev 09)
├scsi 7:x:x:x [Empty]
└scsi 8:x:x:x [Empty]
Other Block Devices
├loop0 0.00k [7:0] Empty/Unknown
├loop1 0.00k [7:1] Empty/Unknown
├loop2 0.00k [7:2] Empty/Unknown
├loop3 0.00k [7:3] Empty/Unknown
├loop4 0.00k [7:4] Empty/Unknown
├loop5 0.00k [7:5] Empty/Unknown
├loop6 0.00k [7:6] Empty/Unknown
└loop7 0.00k [7:7] Empty/Unknown


> >> Before proceeding, please supply more information:
> >> 
> >> for x in /dev/sd[a-p] ; mdadm -E $x ; smartctl -i -A -l scterc $x ;
> >> done
> >> 
> >> Paste the output inline in your response.
> > 
> > 
> > I couldn't get smartctl to work successfully. The version supported
> > on debian squeeze doesn't support aacraid.
> 
> > I tried from a chroot in a debootstrap with a more recent debian
> > version, but only got:
> > 
> > # smartctl --all -d aacraid,0,0,0 /dev/sda
> 
> > smartctl 6.4 2014-10-07 r4002 [x86_64-linux-2.6.32-5-amd64] (local
> > build)
> 
> > Copyright (C) 2002-14, Bruce Allen, Christian Franke,
> > www.smartmontools.org
> > 
> > Smartctl open device: /dev/sda [aacraid_disk_00_00_0] [SCSI/SAT]
> > failed: INQUIRY [SAT]: aacraid result: 0.0 = 22/0
> 
> It's possible the 0,0,0 isn't correct.  The output of lsdrv would help
> with this.
> 
> Also, please use the smartctl options I requested.  '--all' omits the
> scterc information I want to see, and shows a bunch of data I don't need
> to see.  If you want all possible data for your own use, '-x' is the
> correct option.

Yes, I will use this option to filter if I get smartctl to work.

> 
> [trim /]
> 
> It's very important that we get a map of drive serial numbers to current
> device names and the "Device Role" from "mdadm --examine".  As an
> alternative, post the output of "ls -l /dev/disk/by-id/".  This is
> critical information for any future re-create attempts.

lrwxrwxrwx 1 root root  9 Nov 12 10:19 ata-Hitachi_HDP725032GLA360_GEAC34RF2T8G0A -> ../../sdr
lrwxrwxrwx 1 root root 10 Nov 12 10:19 ata-Hitachi_HDP725032GLA360_GEAC34RF2T8G0A-part1 -> ../../sdr1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 ata-Hitachi_HDP725032GLA360_GEAC34RF2T8G0A-part2 -> ../../sdr2
lrwxrwxrwx 1 root root  9 Nov 12 10:19 ata-Hitachi_HDP725032GLA360_GEAC34RF2T8SLA -> ../../sdq
lrwxrwxrwx 1 root root 10 Nov 12 10:19 ata-Hitachi_HDP725032GLA360_GEAC34RF2T8SLA-part1 -> ../../sdq1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 ata-Hitachi_HDP725032GLA360_GEAC34RF2T8SLA-part2 -> ../../sdq2
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-baie-data1 -> ../../dm-3
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-baie-grid5000 -> ../../dm-4
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-baie-home -> ../../dm-2
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-rootvg-dom0--root -> ../../dm-0
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-rootvg-dom0--swap -> ../../dm-1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-rootvg-dom0--tmp -> ../../dm-6
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-rootvg-dom0--var -> ../../dm-5
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-name-rootvg-false_root -> ../../dm-7
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-7krzHXLz487ibYRKTbIZaXzZlz8ju8MM4QRfpRFoJ9EJDP7Nar3SLNj53t7urGbk -> ../../dm-4
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-7krzHXLz487ibYRKTbIZaXzZlz8ju8MMICvtF5UTbncSUMC9f0PyK5zHGmmEa8GD -> ../../dm-2
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-7krzHXLz487ibYRKTbIZaXzZlz8ju8MMkzJJGdeMc0QDg4B1r2hsq5bCnS7Ktk4u -> ../../dm-3
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-oWuOGPc6BtlrebYWwfKkwteqUGfmgRufCqs0FclHYC6O5RNOSEpeRZ3xJ3kXCOG0 -> ../../dm-7
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-oWuOGPc6BtlrebYWwfKkwteqUGfmgRufGm4mzDQtuUTShTEyWgXEo8BXt1d2S4Qu -> ../../dm-1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-oWuOGPc6BtlrebYWwfKkwteqUGfmgRufMGhnq5OTr3pyXgyc2CqDE5ibq9xaOSUf -> ../../dm-5
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-oWuOGPc6BtlrebYWwfKkwteqUGfmgRufOD5FJuWOVLYk7wnRPOvlQOLEb0zffl2X -> ../../dm-0
lrwxrwxrwx 1 root root 10 Nov 12 10:19 dm-uuid-LVM-oWuOGPc6BtlrebYWwfKkwteqUGfmgRufuMkGACbZV71GDBcRVxXnAMf7NkWFWezw -> ../../dm-6
lrwxrwxrwx 1 root root  9 Nov 12 10:19 md-name-ftalc2:0 -> ../../md0
lrwxrwxrwx 1 root root  9 Nov 12 10:19 md-name-ftalc2:1 -> ../../md1
lrwxrwxrwx 1 root root  9 Nov 12 10:19 md-name-ftalc2.nancy.grid5000.fr:2 -> ../../md2
lrwxrwxrwx 1 root root  9 Nov 12 10:19 md-uuid-2d0b91e8:a0b10f4c:3fa285f9:3198a918 -> ../../md2
lrwxrwxrwx 1 root root  9 Nov 12 10:19 md-uuid-791b53cf:48007f45:1dc0ae5f:8cedc958 -> ../../md0
lrwxrwxrwx 1 root root  9 Nov 12 10:19 md-uuid-819ab09a:84026762:9e1f6278:f5bbda51 -> ../../md1
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_0_6F7C0529 -> ../../sda
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_10_897976AC -> ../../sdj
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_11_6DEC1529 -> ../../sdk
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_12_71142529 -> ../../sdl
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_13_14242529 -> ../../sdm
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_14_2D382529 -> ../../sdn
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_15_B4542529 -> ../../sdo
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_1_8E940529 -> ../../sdp
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_2_81A40529 -> ../../sdb
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_3_156214AB -> ../../sdc
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_4_82C40529 -> ../../sdd
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_5_8F341529 -> ../../sde
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_6_5E4C1529 -> ../../sdf
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_7_FF88E4AC -> ../../sdg
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_8_84B41529 -> ../../sdh
lrwxrwxrwx 1 root root  9 Nov 17 10:18 scsi-SAdaptec_LogicalDrv_9_70C41529 -> ../../sdi
lrwxrwxrwx 1 root root  9 Nov 12 10:19 scsi-SATA_Hitachi_HDP7250_GEAC34RF2T8G0A -> ../../sdr
lrwxrwxrwx 1 root root 10 Nov 12 10:19 scsi-SATA_Hitachi_HDP7250_GEAC34RF2T8G0A-part1 -> ../../sdr1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 scsi-SATA_Hitachi_HDP7250_GEAC34RF2T8G0A-part2 -> ../../sdr2
lrwxrwxrwx 1 root root  9 Nov 12 10:19 scsi-SATA_Hitachi_HDP7250_GEAC34RF2T8SLA -> ../../sdq
lrwxrwxrwx 1 root root 10 Nov 12 10:19 scsi-SATA_Hitachi_HDP7250_GEAC34RF2T8SLA-part1 -> ../../sdq1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 scsi-SATA_Hitachi_HDP7250_GEAC34RF2T8SLA-part2 -> ../../sdq2
lrwxrwxrwx 1 root root  9 Nov 12 10:19 wwn-0x5000cca34de737a4 -> ../../sdr
lrwxrwxrwx 1 root root 10 Nov 12 10:19 wwn-0x5000cca34de737a4-part1 -> ../../sdr1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 wwn-0x5000cca34de737a4-part2 -> ../../sdr2
lrwxrwxrwx 1 root root  9 Nov 12 10:19 wwn-0x5000cca34de738cd -> ../../sdq
lrwxrwxrwx 1 root root 10 Nov 12 10:19 wwn-0x5000cca34de738cd-part1 -> ../../sdq1
lrwxrwxrwx 1 root root 10 Nov 12 10:19 wwn-0x5000cca34de738cd-part2 -> ../../sdq2

It seems that the mapping changes at each reboot (two drives that host the operating system had different name across reboots).
Since we re-tighten screws, we didn't reboot though.


> The rest of the information from smartctl is important, and you should
> upgrade your system to a level that supports it, but it can wait for later.
> 
> It might be best to boot into a newer environment strictly for this
> recovery task.  Newer kernels and utilities have more bugfixes and are
> much more robust in emergencies.  I normally use SystemRescueCD [2] for
> emergencies like this.

Ok, if I get stuck on some operations, I'll try with SystemRescueCD.


Regards,

Clément and Marc
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply

* Re: RAID5 hangs in break_stripe_batch_list
From: Martin Svec @ 2015-11-17 13:08 UTC (permalink / raw)
  To: Shaohua Li; +Cc: neilb, linux-raid, target-devel
In-Reply-To: <20151117000409.GA92153@kernel.org>

Dne 17.11.2015 v 1:04 Shaohua Li napsal(a):
> On Mon, Nov 16, 2015 at 12:41:29PM +0100, Martin Svec wrote:
>> Hello,
>>
>> yesterday we had an issue with RAID5 in kernel 4.1.13. The device became unresponsive and RAID
>> module reported the following error:
>>
>> Nov 15 03:44:20 lio-203 kernel: [385878.345689] ------------[ cut here ]------------
>> Nov 15 03:44:20 lio-203 kernel: [385878.345704] WARNING: CPU: 2 PID: 601 at drivers/md/raid5.c:4233
>> break_stripe_batch_list+0x1f4/0x2f0 [raid456]()
>> Nov 15 03:44:20 lio-203 kernel: [385878.345706] Modules linked in: target_core_pscsi
>> target_core_file cpufreq_stats cpufreq_userspace cpufreq_powersave cpufreq_conservative
>> x86_pkg_temp_thermal intel_powerclamp intel_rapl iosf_mbi coretemp kvm_intel raid0 kvm
>> crct10dif_pclmul crc32_pclmul sr_mod iTCO_wdt mgag200 cdrom iTCO_vendor_support ttm dcdbas
>> drm_kms_helper aesni_intel snd_pcm ipmi_devintf drm aes_x86_64 snd_timer lrw gf128mul snd
>> glue_helper joydev evdev soundcore sb_edac i2c_algo_bit ipmi_si ablk_helper 8250_fintek wmi
>> ipmi_msghandler cryptd acpi_power_meter edac_core pcspkr ioatdma mei_me mei lpc_ich dca shpchp
>> mfd_core processor thermal_sys raid456 async_raid6_recov async_memcpy button async_pq async_xor xor
>> async_tx raid6_pq md_mod target_core_iblock iscsi_target_mod target_core_mod configfs autofs4 ext4
>> crc16 mbcache jbd2 dm_mod hid_generic uas usbhid usb_storage hid sg sd_mod bnx2x xhci_pci ehci_pci
>> ptp xhci_hcd ehci_hcd pps_core mdio usbcore megaraid_sas crc32c_generic usb_common crc32c_intel
>> scsi_mod libcrc32c
>> Nov 15 03:44:20 lio-203 kernel: [385878.345748] CPU: 2 PID: 601 Comm: md31_raid5 Not tainted
>> 4.1.13-zoner+ #9
>> Nov 15 03:44:20 lio-203 kernel: [385878.345749] Hardware name: Dell Inc. PowerEdge R730xd/0H21J3,
>> BIOS 1.3.6 06/03/2015
>> Nov 15 03:44:20 lio-203 kernel: [385878.345751]  0000000000000000 ffffffffa03ee3c4 ffffffff81574205
>> 0000000000000000
>> Nov 15 03:44:20 lio-203 kernel: [385878.345753]  ffffffff81072e51 ffff88007501ca50 ffff88007501cad8
>> ffff88006d55d618
>> Nov 15 03:44:20 lio-203 kernel: [385878.345755]  0000000000000000 ffff8802707f83c8 ffffffffa03e4964
>> 0000000000000001
>> Nov 15 03:44:20 lio-203 kernel: [385878.345756] Call Trace:
>> Nov 15 03:44:20 lio-203 kernel: [385878.345764]  [<ffffffff81574205>] ? dump_stack+0x40/0x50
>> Nov 15 03:44:20 lio-203 kernel: [385878.345768]  [<ffffffff81072e51>] ? warn_slowpath_common+0x81/0xb0
>> Nov 15 03:44:20 lio-203 kernel: [385878.345772]  [<ffffffffa03e4964>] ?
>> break_stripe_batch_list+0x1f4/0x2f0 [raid456]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345776]  [<ffffffffa03e86cc>] ? handle_stripe+0x80c/0x2650
>> [raid456]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345781]  [<ffffffff8101d756>] ? native_sched_clock+0x26/0x90
>> Nov 15 03:44:20 lio-203 kernel: [385878.345784]  [<ffffffffa03ea696>] ?
>> handle_active_stripes.isra.46+0x186/0x4e0 [raid456]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345787]  [<ffffffffa03ddab6>] ?
>> raid5_wakeup_stripe_thread+0x96/0x1b0 [raid456]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345790]  [<ffffffffa03eb75d>] ? raid5d+0x49d/0x700 [raid456]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345795]  [<ffffffffa014f166>] ? md_thread+0x126/0x130 [md_mod]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345798]  [<ffffffff810b1e80>] ? wait_woken+0x90/0x90
>> Nov 15 03:44:20 lio-203 kernel: [385878.345801]  [<ffffffffa014f040>] ? find_pers+0x70/0x70 [md_mod]
>> Nov 15 03:44:20 lio-203 kernel: [385878.345805]  [<ffffffff810913d3>] ? kthread+0xd3/0xf0
>> Nov 15 03:44:20 lio-203 kernel: [385878.345807]  [<ffffffff81091300>] ?
>> kthread_create_on_node+0x180/0x180
>> Nov 15 03:44:20 lio-203 kernel: [385878.345811]  [<ffffffff8157a622>] ? ret_from_fork+0x42/0x70
>> Nov 15 03:44:20 lio-203 kernel: [385878.345813]  [<ffffffff81091300>] ?
>> kthread_create_on_node+0x180/0x180
>> Nov 15 03:44:20 lio-203 kernel: [385878.345814] ---[ end trace 298194e8d69e6c62 ]---
>>
>> Unfortunately I'm not able to reproduce the bug, but it seems to be related to high write load. Note
>> that the same issue is also reported here: https://bugzilla.redhat.com/show_bug.cgi?id=1258153 .
>>
>> The setup consists of RAID0 over two RAID5 arrays. Each RAID5 has 6x 960 GB SSD and chunk size 32k.
>> RAID0 has chunk size 160k. Only one of the two RAIDs was affected. After machine reboot, I manually
>> triggered check of both RAID5 arrays and no parity errors were found. Kernel is vanilla stable 4.1.13.
>>
>> Probably there's something wrong with the stripe batching added in 4.1 series? Is there any way to
>> turn the stripe batching off until the bug will be fixed?
> do you have the full dmesg? I'd like to check what triggers the batch break,
> which would be helpful for debugging.

Yes, but I see nothing suspicious before the break_stripe_batch_list warning:

http://pastebin.ca/3258125 ... tail of full dmesg.
http://pastebin.ca/3258121 ... all log entries since last reboot, without the iSCSI
connection/session stuff.

Top-level array is an iblock backend of LIO iSCSI storage with some iSCSI session debug messages
enabled. That's why the log is full of them. However, everything before the RAID5 warning is common
harmless activity of MSFT/ESXi initiators. Subsequent target errors are probably caused by the
unresponsive RAID array and iSCSI session cleanup attempts (Cc'ing target-devel).

The only non-default settings of RAID5 arrays are chunk_size=32k and group_thread_cnt=2.

Thank you,

Martin Svec


^ permalink raw reply

* Re: Reconstruct a RAID 6 that has failed in a non typical manner
From: Phil Turmel @ 2015-11-17 13:25 UTC (permalink / raw)
  To: Marc Pinhede; +Cc: Clement Parisot, linux-raid
In-Reply-To: <402863738.19875205.1447763445794.JavaMail.zimbra@inria.fr>

Good morning Marc, Clément,

On 11/17/2015 07:30 AM, Marc Pinhede wrote:
> Hello,
> 
> Thanks for your answer. Update since our last mail: We saved many
> data thanks to long and boring rsyncs, with countless reboots: during
> rsync, sometime a drive was suddenly considered in 'failed' state by
> the array. The array was still active (with 13 or 12 / 16 disks) but
> 100% of files failed with I/O after that. We were then forced to
> reboot, reassemble the array and restart rsync.

Yes, a miserable task on a large array.  Good to know you saved most (?)
of your data.

> During those long operation, we have been advised to re-tighten our
> storage bay's screws (carri bay). And this is were the magic
> happened. After screwing them back on, no more problem with drive
> considered failed. We only had 4 file copy failures with I/O, but it
> didn't correspond to a drive failing in the array (still working with
> 14/16 drives).

> We can't guarantee than the problem is fixed, but we moved from about
> 10 reboot a day to 5 days of work without problems.

Very good news.  Finding a root cause for a problem greatly raises the
odds future efforts will succeed.

> We now plan to reset and re-introduce one by one the two drive that
> were not recognize by the array, and let the array synchronize,
> rewriting data on those drive. Does it sounds like a good idea to
> you, or do you think it may fails due to some errors?

Since you've identified a real hardware issue that impacted the entire
array, I wouldn't trust it until every drive is thoroughly wiped and
retested.  Use "badblocks -w -p 2" or similar.  Then construct a new
array and restore your saved data.

[trim /]

>> It's very important that we get a map of drive serial numbers to
>> current device names and the "Device Role" from "mdadm --examine".
>> As an alternative, post the output of "ls -l /dev/disk/by-id/".
>> This is critical information for any future re-create attempts.

If you look close at the lsdrv output, you'll see it successfully
acquired drive serial numbers for all drives.  However, they are
reported as Adaptec Logical drives -- these might be generated by the
adaptec firmware, not the real serial numbers.

> It seems that the mapping changes at each reboot (two drives that
> host the operating system had different name across reboots). Since
> we re-tighten screws, we didn't reboot though.

Device names are dependent on device discovery order, which can change
somewhat randomly.  What I've seen with lsdrv is that order doesn't
change within a single controller -- the scsi addresses
{host:bus:target:lun} have consistent bus:target:lun for a given port on
a controller.  I don't have much experience with adaptec devices, so I'd
be curious if it holds true for them.

>> The rest of the information from smartctl is important, and you
>> should upgrade your system to a level that supports it, but it can
>> wait for later.

Consider compiling a local copy of the latest smartctl instead of using
a chroot.  Supply the scsi address shown in lsdrv to the -d aacraid, option.

Regards,

Phil
--
To unsubscribe from this list: send the line "unsubscribe linux-raid" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply

* [PATCH v3 0/7] User namespace mount updates
From: Seth Forshee @ 2015-11-17 16:39 UTC (permalink / raw)
  To: Eric W. Biederman, linux-bcache, dm-devel, linux-raid, linux-mtd,
	linux-fsdevel, linux-security-module, selinux
  Cc: Alexander Viro, Serge Hallyn, Andy Lutomirski, linux-kernel,
	Seth Forshee

Hi Eric,

Here's another update to my patches for user namespace mounts, based on
your for-testing branch. These patches add safeguards necessary to allow
unprivileged mounts and update SELinux and Smack to safely handle
device-backed mounts from unprivileged users.

The v2 posting received very little in the way of feedback, so changes
are minimal. I've made a trivial style change to the Smack changes at
Casey's request, and I've added Stephen's ack for the SELinux changes.

Thanks,
Seth

Andy Lutomirski (1):
  fs: Treat foreign mounts as nosuid

Seth Forshee (6):
  block_dev: Support checking inode permissions in lookup_bdev()
  block_dev: Check permissions towards block device inode when mounting
  mtd: Check permissions towards mtd block device inode when mounting
  selinux: Add support for unprivileged mounts from user namespaces
  userns: Replace in_userns with current_in_userns
  Smack: Handle labels consistently in untrusted mounts

 drivers/md/bcache/super.c      |  2 +-
 drivers/md/dm-table.c          |  2 +-
 drivers/mtd/mtdsuper.c         |  6 +++++-
 fs/block_dev.c                 | 18 +++++++++++++++---
 fs/exec.c                      |  2 +-
 fs/namespace.c                 | 13 +++++++++++++
 fs/quota/quota.c               |  2 +-
 include/linux/fs.h             |  2 +-
 include/linux/mount.h          |  1 +
 include/linux/user_namespace.h |  6 ++----
 kernel/user_namespace.c        |  6 +++---
 security/commoncap.c           |  4 ++--
 security/selinux/hooks.c       | 25 ++++++++++++++++++++++++-
 security/smack/smack_lsm.c     | 29 +++++++++++++++++++----------
 14 files changed, 89 insertions(+), 29 deletions(-)


^ permalink raw reply

* [PATCH v3 1/7] block_dev: Support checking inode permissions in lookup_bdev()
From: Seth Forshee @ 2015-11-17 16:39 UTC (permalink / raw)
  To: Eric W. Biederman, Kent Overstreet, Alasdair Kergon, Mike Snitzer,
	dm-devel, Neil Brown, David Woodhouse, Brian Norris,
	Alexander Viro, Jan Kara, Jeff Layton, J. Bruce Fields
  Cc: Serge Hallyn, Andy Lutomirski, linux-kernel, linux-bcache,
	linux-raid, linux-mtd, linux-fsdevel, linux-security-module,
	selinux, Seth Forshee
In-Reply-To: <1447778351-118699-1-git-send-email-seth.forshee@canonical.com>

When looking up a block device by path no permission check is
done to verify that the user has access to the block device inode
at the specified path. In some cases it may be necessary to
check permissions towards the inode, such as allowing
unprivileged users to mount block devices in user namespaces.

Add an argument to lookup_bdev() to optionally perform this
permission check. A value of 0 skips the permission check and
behaves the same as before. A non-zero value specifies the mask
of access rights required towards the inode at the specified
path. The check is always skipped if the user has CAP_SYS_ADMIN.

All callers of lookup_bdev() currently pass a mask of 0, so this
patch results in no functional change. Subsequent patches will
add permission checks where appropriate.

Signed-off-by: Seth Forshee <seth.forshee@canonical.com>
---
 drivers/md/bcache/super.c |  2 +-
 drivers/md/dm-table.c     |  2 +-
 drivers/mtd/mtdsuper.c    |  2 +-
 fs/block_dev.c            | 13 ++++++++++---
 fs/quota/quota.c          |  2 +-
 include/linux/fs.h        |  2 +-
 6 files changed, 15 insertions(+), 8 deletions(-)

diff --git a/drivers/md/bcache/super.c b/drivers/md/bcache/super.c
index 679a093a3bf6..e8287b0d1dac 100644
--- a/drivers/md/bcache/super.c
+++ b/drivers/md/bcache/super.c
@@ -1926,7 +1926,7 @@ static ssize_t register_bcache(struct kobject *k, struct kobj_attribute *attr,
 				  sb);
 	if (IS_ERR(bdev)) {
 		if (bdev == ERR_PTR(-EBUSY)) {
-			bdev = lookup_bdev(strim(path));
+			bdev = lookup_bdev(strim(path), 0);
 			mutex_lock(&bch_register_lock);
 			if (!IS_ERR(bdev) && bch_is_open(bdev))
 				err = "device already registered";
diff --git a/drivers/md/dm-table.c b/drivers/md/dm-table.c
index e76ed003769e..35bb3ea4cbe2 100644
--- a/drivers/md/dm-table.c
+++ b/drivers/md/dm-table.c
@@ -380,7 +380,7 @@ int dm_get_device(struct dm_target *ti, const char *path, fmode_t mode,
 	BUG_ON(!t);
 
 	/* convert the path to a device */
-	bdev = lookup_bdev(path);
+	bdev = lookup_bdev(path, 0);
 	if (IS_ERR(bdev)) {
 		dev = name_to_dev_t(path);
 		if (!dev)
diff --git a/drivers/mtd/mtdsuper.c b/drivers/mtd/mtdsuper.c
index 20c02a3b7417..b5b60e1af31c 100644
--- a/drivers/mtd/mtdsuper.c
+++ b/drivers/mtd/mtdsuper.c
@@ -176,7 +176,7 @@ struct dentry *mount_mtd(struct file_system_type *fs_type, int flags,
 	/* try the old way - the hack where we allowed users to mount
 	 * /dev/mtdblock$(n) but didn't actually _use_ the blockdev
 	 */
-	bdev = lookup_bdev(dev_name);
+	bdev = lookup_bdev(dev_name, 0);
 	if (IS_ERR(bdev)) {
 		ret = PTR_ERR(bdev);
 		pr_debug("MTDSB: lookup_bdev() returned %d\n", ret);
diff --git a/fs/block_dev.c b/fs/block_dev.c
index 26cee058dc02..f1f0aa7214a3 100644
--- a/fs/block_dev.c
+++ b/fs/block_dev.c
@@ -1396,7 +1396,7 @@ struct block_device *blkdev_get_by_path(const char *path, fmode_t mode,
 	struct block_device *bdev;
 	int err;
 
-	bdev = lookup_bdev(path);
+	bdev = lookup_bdev(path, 0);
 	if (IS_ERR(bdev))
 		return bdev;
 
@@ -1706,12 +1706,14 @@ EXPORT_SYMBOL(ioctl_by_bdev);
 /**
  * lookup_bdev  - lookup a struct block_device by name
  * @pathname:	special file representing the block device
+ * @mask:	rights to check for (%MAY_READ, %MAY_WRITE, %MAY_EXEC)
  *
  * Get a reference to the blockdevice at @pathname in the current
  * namespace if possible and return it.  Return ERR_PTR(error)
- * otherwise.
+ * otherwise.  If @mask is non-zero, check for access rights to the
+ * inode at @pathname.
  */
-struct block_device *lookup_bdev(const char *pathname)
+struct block_device *lookup_bdev(const char *pathname, int mask)
 {
 	struct block_device *bdev;
 	struct inode *inode;
@@ -1726,6 +1728,11 @@ struct block_device *lookup_bdev(const char *pathname)
 		return ERR_PTR(error);
 
 	inode = d_backing_inode(path.dentry);
+	if (mask != 0 && !capable(CAP_SYS_ADMIN)) {
+		error = __inode_permission(inode, mask);
+		if (error)
+			goto fail;
+	}
 	error = -ENOTBLK;
 	if (!S_ISBLK(inode->i_mode))
 		goto fail;
diff --git a/fs/quota/quota.c b/fs/quota/quota.c
index 3746367098fd..a40eaecbd5cc 100644
--- a/fs/quota/quota.c
+++ b/fs/quota/quota.c
@@ -733,7 +733,7 @@ static struct super_block *quotactl_block(const char __user *special, int cmd)
 
 	if (IS_ERR(tmp))
 		return ERR_CAST(tmp);
-	bdev = lookup_bdev(tmp->name);
+	bdev = lookup_bdev(tmp->name, 0);
 	putname(tmp);
 	if (IS_ERR(bdev))
 		return ERR_CAST(bdev);
diff --git a/include/linux/fs.h b/include/linux/fs.h
index 458ee7b213be..cc18dfb0b98e 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -2388,7 +2388,7 @@ static inline void unregister_chrdev(unsigned int major, const char *name)
 #define BLKDEV_MAJOR_HASH_SIZE	255
 extern const char *__bdevname(dev_t, char *buffer);
 extern const char *bdevname(struct block_device *bdev, char *buffer);
-extern struct block_device *lookup_bdev(const char *);
+extern struct block_device *lookup_bdev(const char *, int mask);
 extern void blkdev_show(struct seq_file *,off_t);
 
 #else
-- 
1.9.1

^ permalink raw reply related

* [PATCH v3 2/7] block_dev: Check permissions towards block device inode when mounting
From: Seth Forshee @ 2015-11-17 16:39 UTC (permalink / raw)
  To: Eric W. Biederman, Alexander Viro
  Cc: Serge Hallyn, Andy Lutomirski, linux-kernel, linux-bcache,
	dm-devel, linux-raid, linux-mtd, linux-fsdevel,
	linux-security-module, selinux, Seth Forshee
In-Reply-To: <1447778351-118699-1-git-send-email-seth.forshee@canonical.com>

Unprivileged users should not be able to mount block devices when
they lack sufficient privileges towards the block device inode.
Update blkdev_get_by_path() to validate that the user has the
required access to the inode at the specified path. The check
will be skipped for CAP_SYS_ADMIN, so privileged mounts will
continue working as before.

Signed-off-by: Seth Forshee <seth.forshee@canonical.com>
---
 fs/block_dev.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/fs/block_dev.c b/fs/block_dev.c
index f1f0aa7214a3..54d94cd64577 100644
--- a/fs/block_dev.c
+++ b/fs/block_dev.c
@@ -1394,9 +1394,14 @@ struct block_device *blkdev_get_by_path(const char *path, fmode_t mode,
 					void *holder)
 {
 	struct block_device *bdev;
+	int perm = 0;
 	int err;
 
-	bdev = lookup_bdev(path, 0);
+	if (mode & FMODE_READ)
+		perm |= MAY_READ;
+	if (mode & FMODE_WRITE)
+		perm |= MAY_WRITE;
+	bdev = lookup_bdev(path, perm);
 	if (IS_ERR(bdev))
 		return bdev;
 
-- 
1.9.1

^ permalink raw reply related

* [PATCH v3 3/7] mtd: Check permissions towards mtd block device inode when mounting
From: Seth Forshee @ 2015-11-17 16:39 UTC (permalink / raw)
  To: Eric W. Biederman, David Woodhouse, Brian Norris
  Cc: Alexander Viro, Serge Hallyn, Andy Lutomirski, linux-kernel,
	linux-bcache, dm-devel, linux-raid, linux-mtd, linux-fsdevel,
	linux-security-module, selinux, Seth Forshee
In-Reply-To: <1447778351-118699-1-git-send-email-seth.forshee@canonical.com>

Unprivileged users should not be able to mount mtd block devices
when they lack sufficient privileges towards the block device
inode.  Update mount_mtd() to validate that the user has the
required access to the inode at the specified path. The check
will be skipped for CAP_SYS_ADMIN, so privileged mounts will
continue working as before.

Signed-off-by: Seth Forshee <seth.forshee@canonical.com>
---
 drivers/mtd/mtdsuper.c | 6 +++++-
 1 file changed, 5 insertions(+), 1 deletion(-)

diff --git a/drivers/mtd/mtdsuper.c b/drivers/mtd/mtdsuper.c
index b5b60e1af31c..5d7e7705fed8 100644
--- a/drivers/mtd/mtdsuper.c
+++ b/drivers/mtd/mtdsuper.c
@@ -125,6 +125,7 @@ struct dentry *mount_mtd(struct file_system_type *fs_type, int flags,
 #ifdef CONFIG_BLOCK
 	struct block_device *bdev;
 	int ret, major;
+	int perm;
 #endif
 	int mtdnr;
 
@@ -176,7 +177,10 @@ struct dentry *mount_mtd(struct file_system_type *fs_type, int flags,
 	/* try the old way - the hack where we allowed users to mount
 	 * /dev/mtdblock$(n) but didn't actually _use_ the blockdev
 	 */
-	bdev = lookup_bdev(dev_name, 0);
+	perm = MAY_READ;
+	if (!(flags & MS_RDONLY))
+		perm |= MAY_WRITE;
+	bdev = lookup_bdev(dev_name, perm);
 	if (IS_ERR(bdev)) {
 		ret = PTR_ERR(bdev);
 		pr_debug("MTDSB: lookup_bdev() returned %d\n", ret);
-- 
1.9.1

^ permalink raw reply related


This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox