CEPH filesystem development
 help / color / mirror / Atom feed
* disk failure prediction
@ 2015-02-18 23:20 Sage Weil
  2015-02-19  8:56 ` [ceph-calamari] " John Spray
  0 siblings, 1 reply; 4+ messages in thread
From: Sage Weil @ 2015-02-18 23:20 UTC (permalink / raw)
  To: ceph-devel, ceph-calamari

Interesting paper at FAST:

	https://www.usenix.org/system/files/conference/fast15/fast15-paper-ma.pdf

Short version: reallocated sectors correllates with impending disk 
failures (this sounds like what Sandon has been telling us for ages) and 
by preemptively replacing disks with impending failures reduced EMC's rate 
of triple-failures by 80%, and looking at the joint failure probability 
within each raid set reduces the failure rate by 98%.  We wouldn't see 
quite the same results since our "raid sets" are effectively entire pools, 
but this seems like a strong case for adding smart monitoring to the osds 
or to calamari already and doing some preemptive disk replacement.

sage

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2015-02-19 15:59 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2015-02-18 23:20 disk failure prediction Sage Weil
2015-02-19  8:56 ` [ceph-calamari] " John Spray
2015-02-19 14:58   ` Sage Weil
2015-02-19 15:59     ` Gregory Meno

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox