Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Back during the dot-com era we lost a production array to the infamous "DeathStar" drives.

The system reported a failure, so we scheduled the drive to be replaced and brought up the hot spare and started the parity resync process. A little while later there was another drive failure and we told the data center folks to tell the tech to hurry up. While the tech was headed to our cage, there was a third drive failure and the array was toast. We were able to restore from backup, but the data was a day old.

Lessons were: Mix drives from different production batches (we couldn't mix manufacturers because of the leasing contract). Have a backup that you can restore from. Parity resync operations while the array is in use will put more stress on the drives than production use alone will, and will kill any (remaining) weak drives.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: