Linux-admin Development Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Glynn Clements <glynn@gclements.plus.com>
To: "Jevos, Peter" <Peter.Jevos@Oriflame-SW.Com>
Cc: linux-admin@vger.kernel.org
Subject: Re: Computer suddenly failed
Date: Tue, 28 Feb 2006 21:19:02 +0000	[thread overview]
Message-ID: <17412.48710.700491.633046@cerise.gclements.plus.com> (raw)
In-Reply-To: <6639BA20B265594590E97DD72A8EE0359DA0C1@exosw.osw.ori.local>


Jevos, Peter wrote:

> I'd like to ask you about strange problem. I hope I chose a correct
> mailing list
> I have 2 IDE disks in RAID 1 with Reiserfs.
> Once I noticed message in the log:
> 
> hde: dma_timer_expiry: dma status == 0x20
> hde: timeout waiting for DMA
> PDC202XX: Primary channel reset.
> hde: timeout waiting for DMA
> hde: (__ide_dma_test_irq) called while not waiting
> hde: status timeout: status=0xd0 { Busy }
>  PDC202XX: Primary channel reset.
> hde: drive not ready for command
> ide2: reset: success
> hde: status error: status=0x58 { DriveReady SeekComplete DataRequest }
> hde: drive not ready for command
> hde: status error: status=0x58 { DriveReady SeekComplete DataRequest }

Your drive has died.

> DMA on hde was turned off so I turned it on again. Than I tried to made
> files backup on the hde, but when I ran tar computer didn't response
> even for sysrq, no log was written. I had to made a hard restart. It
> repeats for 4 times when I tried to did something with files on the hde.
> Now I'm afraid to do anything on that machine,unfortunately it is
> production server.

Once you have replaced the drive:

1. Ensure that the drives are being cooled. Modern (i.e. large) hard
drives tend to run quite hot. In the absence of sufficient airflow,
they can easily exceed their maximum operating temperature (typically
55C). This is more of an issue with larger drives, and with multiple
drives in adjacent drive bays. In my experience, Maxtor drives tend to
run hotter than similar drives from other vendors.

2. Run a temperature-monitoring utility such as hddtemp, and ensure
that it will notify support staff if the temperature gets too high. If
the location isn't staffed 24/7, ensure that it will shut down the
system in the event that the temperature exceeds the drives' operating
limit.

[I know of a case where a cooling fan in a file server failed
overnight, and the staff turned up the following morning to find that
all 4 drives had failed after reaching temperatures of up to 63C.]

-- 
Glynn Clements <glynn@gclements.plus.com>

  parent reply	other threads:[~2006-02-28 21:19 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2006-02-28 11:05 Computer suddenly failed Jevos, Peter
2006-02-28 11:26 ` urgrue
     [not found] ` <6639BA20B265594590E97DD72A8EE0359DA0C1@exosw.osw.ori.local >
2006-02-28 11:51   ` Adrian C.
2006-02-28 20:57     ` Glynn Clements
2006-02-28 21:19 ` Glynn Clements [this message]
  -- strict thread matches above, loose matches on Subject: below --
2006-02-28 12:29 Jevos, Peter

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=17412.48710.700491.633046@cerise.gclements.plus.com \
    --to=glynn@gclements.plus.com \
    --cc=Peter.Jevos@Oriflame-SW.Com \
    --cc=linux-admin@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox