Linux RAID subsystem development
 help / color / mirror / Atom feed
* Fedora 20 RAID 6 errors on rebuild / check / repair
@ 2014-07-24  2:29 George Rapp
  2014-07-24  7:33 ` Kay Diederichs
  0 siblings, 1 reply; 5+ messages in thread
From: George Rapp @ 2014-07-24  2:29 UTC (permalink / raw)
  To: linux-raid

Hi -

I have a Fedora 20 media server / MythTV backend utilizing a HighPoint
RocketRAID 2720SGL controller (Amazon product link:
http://is.gd/yqo2i1). The server performs fine under normal (minimal)
read-write operations, but during any high-I/O operations (rebuild
after mdadm --add, RAID check initiated by "echo check >
/sys/block/md6/md/sync_action" or "echo repair > ..."), I get sporadic
errors and poor performance on my RAID 6 array, /dev/md6.

Wondering if there is anything I can tweak to make my configuration
more stable. The inability to check or repair this RAID device has me
nervous.

The problems seem to start when I see the following error message in
/var/log/syslog:

> Jul 22 21:23:37 backend3 kernel: [95876.375990] ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
> Jul 22 21:23:37 backend3 kernel: [95876.376153] ata5.00: failed command: READ DMA
> Jul 22 21:23:37 backend3 kernel: [95876.376284] ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
> Jul 22 21:23:37 backend3 kernel: [95876.376284]          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
> Jul 22 21:23:37 backend3 kernel: [95876.376750] ata5.00: status: { DRDY }
> Jul 22 21:23:37 backend3 kernel: [95876.376874] ata5: hard resetting link
> Jul 22 21:23:37 backend3 kernel: ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
> Jul 22 21:23:37 backend3 kernel: ata5.00: failed command: READ DMA
> Jul 22 21:23:37 backend3 kernel: ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
>          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
> Jul 22 21:23:37 backend3 kernel: ata5.00: status: { DRDY }
> Jul 22 21:23:37 backend3 kernel: ata5: hard resetting link
> Jul 22 21:23:40 backend3 kernel: [95878.742281] ata5.00: configured for UDMA/133
> Jul 22 21:23:40 backend3 kernel: [95878.742413] ata5.00: device reported invalid CHS sector 0
> Jul 22 21:23:40 backend3 kernel: [95878.742542] ata5: EH complete
> Jul 22 21:23:40 backend3 kernel: ata5.00: configured for UDMA/133
> Jul 22 21:23:40 backend3 kernel: ata5.00: device reported invalid CHS sector 0
> Jul 22 21:23:40 backend3 kernel: ata5: EH complete


I thought the problem might be caused by NCQ being enabled -- previous
iterations of this error included the string 'ncq', like this:

> ata7.00: cmd 60/00:00:68:4b:75/03:00:04:00:00/40 tag 0 ncq 393216 in
>          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)


so I disabled NCQ by adding "libata.force=noncq" to my kernel boot
parameters. However, it didn't help, as I still get the "...frozen"
errors. (I have young children, so any error message that includes the
word "Frozen" makes me twitchy ... 8^)

Right now, I'm attempting to rebuild the degraded RAID 6 array after
swapping out a disk that was getting an increasing number of these
errors:

> Device: /dev/sdf [SAT], 35 Currently unreadable (pending) sectors


I started the rebuild on Monday night in single-user mode via

> # mdadm --manage /dev/md6 --add /dev/sdf1


(my other partitions are /dev/sd[bcde]4, but I only created a single
partition on the new disk, to see if reliability would be better by
placing my RAID partition on the first partition rather than the last)

At first, the rebuild was supposed to take 4.3 days. I Googled around
and found a couple of speed optimization techniques, which I applied:

> # sysctl -w dev.raid.speed_limit_max=100000
>
> # cd /sys/block/md6/md
> # echo 16384 > stripe_cache_size


This initially sped up the resync speed to 68-70000K/sec, until I hit
the first "exception Emask" error like the one I described above --
now the speed has dropped to 30K/sec, and the rebuild is scheduled to
last 439 more days! I don't know if I should just mark the new device
as failed and stop the sync, or let it keep grinding and hope it
speeds up.

Any pointers or tips appreciated. I've been running Linux software
RAID for 4-5 years, but this is the first time I've experienced this
kind of trouble.

More data on my system:


[root@backend3 gwr]# uname -a
Linux backend3 3.14.4-200.fc20.i686+PAE #1 SMP Tue May 13 14:03:12 UTC
2014 i686 i686 i386 GNU/Linux

[root@backend3 gwr]# mdadm --version
mdadm - v3.3 - 3rd September 2013

[root@backend3 log]# mdadm --detail /dev/md6
/dev/md6:
        Version : 1.2
  Creation Time : Sun Apr 24 17:31:27 2011
     Raid Level : raid6
     Array Size : 5756723712 (5490.04 GiB 5894.89 GB)
  Used Dev Size : 1918907904 (1830.01 GiB 1964.96 GB)
   Raid Devices : 5
  Total Devices : 5
    Persistence : Superblock is persistent

  Intent Bitmap : Internal

    Update Time : Wed Jul 23 22:15:05 2014
          State : active, degraded, recovering
 Active Devices : 4
Working Devices : 5
 Failed Devices : 0
  Spare Devices : 1

         Layout : left-symmetric
     Chunk Size : 512K

 Rebuild Status : 39% complete

           Name : backend3:md4
           UUID : 894bc20e:b9479ac9:7bfce54f:0ac12dd9
         Events : 1659380

    Number   Major   Minor   RaidDevice State
       0       8       36        0      active sync   /dev/sdc4
       1       8       20        1      active sync   /dev/sdb4
       5       8       68        2      active sync   /dev/sde4
       4       8       52        3      active sync   /dev/sdd4
       6       8       81        4      spare rebuilding   /dev/sdf1

-- 
George Rapp  (Pataskala, OH) Home: george.rapp -- at -- gmail.com
Work: george.rapp -- at -- hp.com (or) george.rapp.ctr -- at -- dfas.mil

A wise and frugal government, which shall restrain men from injuring
one another, which shall leave them otherwise free to regulate their
own pursuits of industry and improvement, and shall not take from the
mouth of labor the bread it has earned. This is the sum of
good government... - Thomas Jefferson, First Inaugural Address

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: Fedora 20 RAID 6 errors on rebuild / check / repair
  2014-07-24  2:29 Fedora 20 RAID 6 errors on rebuild / check / repair George Rapp
@ 2014-07-24  7:33 ` Kay Diederichs
  2014-07-24  9:45   ` Wilson, Jonathan
  2014-07-25 16:57   ` George Rapp
  0 siblings, 2 replies; 5+ messages in thread
From: Kay Diederichs @ 2014-07-24  7:33 UTC (permalink / raw)
  To: linux-raid

On 07/24/2014 04:29 AM, George Rapp wrote:
> Hi -
> 
> I have a Fedora 20 media server / MythTV backend utilizing a HighPoint
> RocketRAID 2720SGL controller (Amazon product link:
> http://is.gd/yqo2i1). The server performs fine under normal (minimal)
> read-write operations, but during any high-I/O operations (rebuild
> after mdadm --add, RAID check initiated by "echo check >
> /sys/block/md6/md/sync_action" or "echo repair > ..."), I get sporadic
> errors and poor performance on my RAID 6 array, /dev/md6.
> 
> Wondering if there is anything I can tweak to make my configuration
> more stable. The inability to check or repair this RAID device has me
> nervous.
> 
> The problems seem to start when I see the following error message in
> /var/log/syslog:
> 
>> Jul 22 21:23:37 backend3 kernel: [95876.375990] ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
>> Jul 22 21:23:37 backend3 kernel: [95876.376153] ata5.00: failed command: READ DMA
>> Jul 22 21:23:37 backend3 kernel: [95876.376284] ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
>> Jul 22 21:23:37 backend3 kernel: [95876.376284]          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
>> Jul 22 21:23:37 backend3 kernel: [95876.376750] ata5.00: status: { DRDY }
>> Jul 22 21:23:37 backend3 kernel: [95876.376874] ata5: hard resetting link
>> Jul 22 21:23:37 backend3 kernel: ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
>> Jul 22 21:23:37 backend3 kernel: ata5.00: failed command: READ DMA
>> Jul 22 21:23:37 backend3 kernel: ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
>>          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
>> Jul 22 21:23:37 backend3 kernel: ata5.00: status: { DRDY }
>> Jul 22 21:23:37 backend3 kernel: ata5: hard resetting link
>> Jul 22 21:23:40 backend3 kernel: [95878.742281] ata5.00: configured for UDMA/133
>> Jul 22 21:23:40 backend3 kernel: [95878.742413] ata5.00: device reported invalid CHS sector 0
>> Jul 22 21:23:40 backend3 kernel: [95878.742542] ata5: EH complete
>> Jul 22 21:23:40 backend3 kernel: ata5.00: configured for UDMA/133
>> Jul 22 21:23:40 backend3 kernel: ata5.00: device reported invalid CHS sector 0
>> Jul 22 21:23:40 backend3 kernel: ata5: EH complete
> 
> 
> I thought the problem might be caused by NCQ being enabled -- previous
> iterations of this error included the string 'ncq', like this:
> 
>> ata7.00: cmd 60/00:00:68:4b:75/03:00:04:00:00/40 tag 0 ncq 393216 in
>>          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
> 
> 
> so I disabled NCQ by adding "libata.force=noncq" to my kernel boot
> parameters. However, it didn't help, as I still get the "...frozen"
> errors. (I have young children, so any error message that includes the
> word "Frozen" makes me twitchy ... 8^)
> 
> Right now, I'm attempting to rebuild the degraded RAID 6 array after
> swapping out a disk that was getting an increasing number of these
> errors:
> 
>> Device: /dev/sdf [SAT], 35 Currently unreadable (pending) sectors
> 
> 
> I started the rebuild on Monday night in single-user mode via
> 
>> # mdadm --manage /dev/md6 --add /dev/sdf1
> 
> 
> (my other partitions are /dev/sd[bcde]4, but I only created a single
> partition on the new disk, to see if reliability would be better by
> placing my RAID partition on the first partition rather than the last)
> 
> At first, the rebuild was supposed to take 4.3 days. I Googled around
> and found a couple of speed optimization techniques, which I applied:
> 
>> # sysctl -w dev.raid.speed_limit_max=100000
>>
>> # cd /sys/block/md6/md
>> # echo 16384 > stripe_cache_size
> 
> 
> This initially sped up the resync speed to 68-70000K/sec, until I hit
> the first "exception Emask" error like the one I described above --
> now the speed has dropped to 30K/sec, and the rebuild is scheduled to
> last 439 more days! I don't know if I should just mark the new device
> as failed and stop the sync, or let it keep grinding and hope it
> speeds up.
> 
> Any pointers or tips appreciated. I've been running Linux software
> RAID for 4-5 years, but this is the first time I've experienced this
> kind of trouble.
> 
> More data on my system:
> 
> 
> [root@backend3 gwr]# uname -a
> Linux backend3 3.14.4-200.fc20.i686+PAE #1 SMP Tue May 13 14:03:12 UTC
> 2014 i686 i686 i386 GNU/Linux
> 
> [root@backend3 gwr]# mdadm --version
> mdadm - v3.3 - 3rd September 2013
> 
> [root@backend3 log]# mdadm --detail /dev/md6
> /dev/md6:
>         Version : 1.2
>   Creation Time : Sun Apr 24 17:31:27 2011
>      Raid Level : raid6
>      Array Size : 5756723712 (5490.04 GiB 5894.89 GB)
>   Used Dev Size : 1918907904 (1830.01 GiB 1964.96 GB)
>    Raid Devices : 5
>   Total Devices : 5
>     Persistence : Superblock is persistent
> 
>   Intent Bitmap : Internal
> 
>     Update Time : Wed Jul 23 22:15:05 2014
>           State : active, degraded, recovering
>  Active Devices : 4
> Working Devices : 5
>  Failed Devices : 0
>   Spare Devices : 1
> 
>          Layout : left-symmetric
>      Chunk Size : 512K
> 
>  Rebuild Status : 39% complete
> 
>            Name : backend3:md4
>            UUID : 894bc20e:b9479ac9:7bfce54f:0ac12dd9
>          Events : 1659380
> 
>     Number   Major   Minor   RaidDevice State
>        0       8       36        0      active sync   /dev/sdc4
>        1       8       20        1      active sync   /dev/sdb4
>        5       8       68        2      active sync   /dev/sde4
>        4       8       52        3      active sync   /dev/sdd4
>        6       8       81        4      spare rebuilding   /dev/sdf1
> 

George,

this is not necessarily a RAID problem. Can you exclude the possibility
that one or more of the disks have a hardware problem, like the one you
replaced which showed
>> Device: /dev/sdf [SAT], 35 Currently unreadable (pending) sectors
Hardware problems would explain the problems you have.

What does smartctl report about your disks, in particular:
Offline_Uncorrectable
Current_Pending_Sector
Reallocated_Sector_Ct

And, is it always ATA 5.00 that is mentioned in syslog? dmesg and the
"lsdrv" script (google for it) are useful in diagnosing this.

HTH,
Kay



^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: Fedora 20 RAID 6 errors on rebuild / check / repair
  2014-07-24  7:33 ` Kay Diederichs
@ 2014-07-24  9:45   ` Wilson, Jonathan
  2014-07-24 19:22     ` George Rapp
  2014-07-25 16:57   ` George Rapp
  1 sibling, 1 reply; 5+ messages in thread
From: Wilson, Jonathan @ 2014-07-24  9:45 UTC (permalink / raw)
  To: Kay Diederichs; +Cc: linux-raid

On Thu, 2014-07-24 at 09:33 +0200, Kay Diederichs wrote:
> On 07/24/2014 04:29 AM, George Rapp wrote:
> > Hi -
> > 
> > I have a Fedora 20 media server / MythTV backend utilizing a HighPoint
> > RocketRAID 2720SGL controller (Amazon product link:
> > http://is.gd/yqo2i1). The server performs fine under normal (minimal)
> > read-write operations, but during any high-I/O operations (rebuild
> > after mdadm --add, RAID check initiated by "echo check >
> > /sys/block/md6/md/sync_action" or "echo repair > ..."), I get sporadic
> > errors and poor performance on my RAID 6 array, /dev/md6.
> > 
> > Wondering if there is anything I can tweak to make my configuration
> > more stable. The inability to check or repair this RAID device has me
> > nervous.
> > 
> > The problems seem to start when I see the following error message in
> > /var/log/syslog:
> > 
> >> Jul 22 21:23:37 backend3 kernel: [95876.375990] ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
> >> Jul 22 21:23:37 backend3 kernel: [95876.376153] ata5.00: failed command: READ DMA
> >> Jul 22 21:23:37 backend3 kernel: [95876.376284] ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
> >> Jul 22 21:23:37 backend3 kernel: [95876.376284]          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
> >> Jul 22 21:23:37 backend3 kernel: [95876.376750] ata5.00: status: { DRDY }
> >> Jul 22 21:23:37 backend3 kernel: [95876.376874] ata5: hard resetting link
> >> Jul 22 21:23:37 backend3 kernel: ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
> >> Jul 22 21:23:37 backend3 kernel: ata5.00: failed command: READ DMA
> >> Jul 22 21:23:37 backend3 kernel: ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
> >>          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
> >> Jul 22 21:23:37 backend3 kernel: ata5.00: status: { DRDY }
> >> Jul 22 21:23:37 backend3 kernel: ata5: hard resetting link
> >> Jul 22 21:23:40 backend3 kernel: [95878.742281] ata5.00: configured for UDMA/133
> >> Jul 22 21:23:40 backend3 kernel: [95878.742413] ata5.00: device reported invalid CHS sector 0
> >> Jul 22 21:23:40 backend3 kernel: [95878.742542] ata5: EH complete
> >> Jul 22 21:23:40 backend3 kernel: ata5.00: configured for UDMA/133
> >> Jul 22 21:23:40 backend3 kernel: ata5.00: device reported invalid CHS sector 0
> >> Jul 22 21:23:40 backend3 kernel: ata5: EH complete
> > 

The above tends to point to a hardware problem with any of the
following.. disk, cable, controller.

My own experience of such messages, they where always caused by
connection problems in the cables with one being "broken" in a similar
way to how a pair of head phones cut out until the cable is "wobbled"
near the jack plug.

Basically a broken wire in the sata cable that works "most of the time"
but under load fails. It was very badly "kinked" near the plug due to
bad case design not allowing much room between the side panel and the
back of the drive and over time and multiple side panel removals moving
the cable to different drives had degraded the integrity.

The second time I had the above was caused by a loose socket on an add
in card (really cheap one), the vibrations from the washing machine spin
dry sequence would cause it to error if under high load (working out
that two unrelated events had to occur at the same time took a while,
especially as the drive in the second sata socket never had any issues);
it was eventually resolved by using a small piece of selotape which
lifted one side of the cable connector causing a tighter fit on the
contact side... a true "bodge-it and scarper" fix that would make an
engineer proud ;-)




^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: Fedora 20 RAID 6 errors on rebuild / check / repair
  2014-07-24  9:45   ` Wilson, Jonathan
@ 2014-07-24 19:22     ` George Rapp
  0 siblings, 0 replies; 5+ messages in thread
From: George Rapp @ 2014-07-24 19:22 UTC (permalink / raw)
  To: Wilson, Jonathan; +Cc: Kay Diederichs, linux-raid

On Thu, Jul 24, 2014 at 5:45 AM, Wilson, Jonathan
<piercing_male@hotmail.com> wrote:
> On Thu, 2014-07-24 at 09:33 +0200, Kay Diederichs wrote:
>> On 07/24/2014 04:29 AM, George Rapp wrote:
>> > [snip]
>> >
>> >> Jul 22 21:23:37 backend3 kernel: [95876.375990] ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
>> >> Jul 22 21:23:37 backend3 kernel: [95876.376153] ata5.00: failed command: READ DMA
>> >> Jul 22 21:23:37 backend3 kernel: [95876.376284] ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
>> >> Jul 22 21:23:37 backend3 kernel: [95876.376284]          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
>> >> Jul 22 21:23:37 backend3 kernel: [95876.376750] ata5.00: status: { DRDY }
>> >> Jul 22 21:23:37 backend3 kernel: [95876.376874] ata5: hard resetting link
>> >> Jul 22 21:23:37 backend3 kernel: ata5.00: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
>> >> Jul 22 21:23:37 backend3 kernel: ata5.00: failed command: READ DMA
>> >> Jul 22 21:23:37 backend3 kernel: ata5.00: cmd c8/00:08:40:11:81/00:00:00:00:00/e3 tag 11 dma 4096 in
>> >>          res 40/00:00:00:4f:c2/00:00:00:00:00/40 Emask 0x4 (timeout)
>> >> Jul 22 21:23:37 backend3 kernel: ata5.00: status: { DRDY }
>> >> Jul 22 21:23:37 backend3 kernel: ata5: hard resetting link
>> >> Jul 22 21:23:40 backend3 kernel: [95878.742281] ata5.00: configured for UDMA/133
>> >> Jul 22 21:23:40 backend3 kernel: [95878.742413] ata5.00: device reported invalid CHS sector 0
>> >> Jul 22 21:23:40 backend3 kernel: [95878.742542] ata5: EH complete
>> >> Jul 22 21:23:40 backend3 kernel: ata5.00: configured for UDMA/133
>> >> Jul 22 21:23:40 backend3 kernel: ata5.00: device reported invalid CHS sector 0
>> >> Jul 22 21:23:40 backend3 kernel: ata5: EH complete
>> >
>
> The above tends to point to a hardware problem with any of the
> following.. disk, cable, controller.
>
> My own experience of such messages, they where always caused by
> connection problems in the cables with one being "broken" in a similar
> way to how a pair of head phones cut out until the cable is "wobbled"
> near the jack plug.
>
> Basically a broken wire in the sata cable that works "most of the time"
> but under load fails. It was very badly "kinked" near the plug due to
> bad case design not allowing much room between the side panel and the
> back of the drive and over time and multiple side panel removals moving
> the cable to different drives had degraded the integrity.

Jonathan -

Thanks for your reply. I am planning on moving this system to another
case that (1) is larger and (2) has some better ability to attenuate
vibrations. I might also replace my cables.

> The second time I had the above was caused by a loose socket on an add
> in card (really cheap one), the vibrations from the washing machine spin
> dry sequence would cause it to error if under high load (working out
> that two unrelated events had to occur at the same time took a while,
> especially as the drive in the second sata socket never had any issues);
> it was eventually resolved by using a small piece of selotape which
> lifted one side of the cable connector causing a tighter fit on the
> contact side... a true "bodge-it and scarper" fix that would make an
> engineer proud ;-)

The last three tools used on this server have been epoxy, a pop rivet
gun, and a sheet metal bending bar. I understand your solution (what
we in the US would call "redneck engineering") perfectly... 8^)

Thanks again.
-- 
George Rapp  (Pataskala, OH) Home: george.rapp -- at -- gmail.com
Work: george.rapp -- at -- hp.com (or) george.rapp.ctr -- at -- dfas.mil

A wise and frugal government, which shall restrain men from injuring
one another, which shall leave them otherwise free to regulate their
own pursuits of industry and improvement, and shall not take from the
mouth of labor the bread it has earned. This is the sum of
good government... - Thomas Jefferson, First Inaugural Address

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: Fedora 20 RAID 6 errors on rebuild / check / repair
  2014-07-24  7:33 ` Kay Diederichs
  2014-07-24  9:45   ` Wilson, Jonathan
@ 2014-07-25 16:57   ` George Rapp
  1 sibling, 0 replies; 5+ messages in thread
From: George Rapp @ 2014-07-25 16:57 UTC (permalink / raw)
  To: Kay Diederichs; +Cc: linux-raid

On Thu, Jul 24, 2014 at 3:33 AM, Kay Diederichs
<kay.diederichs@uni-konstanz.de> wrote:
> On 07/24/2014 04:29 AM, George Rapp wrote:
>> Hi -
>> [big snip]
>>

Kay

Thanks for your feedback. I'll try to address the details of what you
asked below.

> this is not necessarily a RAID problem. Can you exclude the possibility
> that one or more of the disks have a hardware problem, like the one you
> replaced which showed
>>> Device: /dev/sdf [SAT], 35 Currently unreadable (pending) sectors
> Hardware problems would explain the problems you have.
>
> What does smartctl report about your disks, in particular:
> Offline_Uncorrectable
> Current_Pending_Sector
> Reallocated_Sector_Ct

I don't think any of my current disks are showing signs of hardware
problems, looking for the values you suggested:

[root@backend3 gwr]# smartctl -a /dev/sdb | grep -E
"Uncorrectable|Pending|Reallocated"
  5 Reallocated_Sector_Ct   0x0033   100   100   036    Pre-fail
Always       -       0
197 Current_Pending_Sector  0x0012   100   100   000    Old_age
Always       -       0
198 Offline_Uncorrectable   0x0010   100   100   000    Old_age
Offline      -       0

[root@backend3 gwr]# smartctl -a /dev/sdd | grep -E
"Uncorrectable|Pending|Reallocated"
  5 Reallocated_Sector_Ct   0x0033   100   100   036    Pre-fail
Always       -       0
197 Current_Pending_Sector  0x0012   100   100   000    Old_age
Always       -       8
198 Offline_Uncorrectable   0x0010   100   100   000    Old_age
Offline      -       8

[root@backend3 gwr]# smartctl -a /dev/sde | grep -E
"Uncorrectable|Pending|Reallocated"
  5 Reallocated_Sector_Ct   0x0033   200   200   140    Pre-fail
Always       -       0
196 Reallocated_Event_Count 0x0032   200   200   000    Old_age
Always       -       0
197 Current_Pending_Sector  0x0032   200   200   000    Old_age
Always       -       0
198 Offline_Uncorrectable   0x0030   100   253   000    Old_age
Offline      -       0

[root@backend3 gwr]# smartctl -a /dev/sdf | grep -E
"Uncorrectable|Pending|Reallocated"
  5 Reallocated_Sector_Ct   0x0033   100   100   036    Pre-fail
Always       -       0
197 Current_Pending_Sector  0x0012   100   100   000    Old_age
Always       -       0
198 Offline_Uncorrectable   0x0010   100   100   000    Old_age
Offline      -       0

I can't get smartctl to tell me anything about /dev/sdc at the moment

> And, is it always ATA 5.00 that is mentioned in syslog? dmesg and the
> "lsdrv" script (google for it) are useful in diagnosing this.

No, all five disks in the array (ATA 5.00 through 9.00) are mentioned
at times in syslog. There doesn't appear to be any rhyme nor reason to
it.


-- 
George Rapp  (Pataskala, OH) Home: george.rapp -- at -- gmail.com
Work: george.rapp -- at -- hp.com (or) george.rapp.ctr -- at -- dfas.mil

A wise and frugal government, which shall restrain men from injuring
one another, which shall leave them otherwise free to regulate their
own pursuits of industry and improvement, and shall not take from the
mouth of labor the bread it has earned. This is the sum of
good government... - Thomas Jefferson, First Inaugural Address

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2014-07-25 16:57 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2014-07-24  2:29 Fedora 20 RAID 6 errors on rebuild / check / repair George Rapp
2014-07-24  7:33 ` Kay Diederichs
2014-07-24  9:45   ` Wilson, Jonathan
2014-07-24 19:22     ` George Rapp
2014-07-25 16:57   ` George Rapp

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox