A RAID array that refuses to rebuild — or that fails partway through and drops offline — is one of the most nerve-wracking things a server or NAS can do. The instinct is to try again, swap another disk, or hit “rebuild” once more. Please don’t. A stalled rebuild is a warning, and forcing it is how recoverable arrays become permanently lost. Here’s why they fail, and the safe way through.
Rebuilds fail because a second disk is weak, the controller is confused, or the disk order is wrong. Forcing or repeating the rebuild stresses the surviving disks and can overwrite the data needed to recover the array. Stop first.
A rebuild that won’t complete usually has a specific cause. The commonest is a second weak disk: RAID disks are typically bought together and wear together, so when one fails, another is often close behind — and the intense reading a rebuild demands is exactly what tips it over. Other causes include a controller or configuration problem (the array loses its configuration, or disks are detected in the wrong order), a disk that’s been replaced with an incompatible one, or simple unreadable sectors on an otherwise working disk that halt the parity calculation. None of these are fixed by trying again — and trying again is what does the damage.
Here’s the mechanism that catches people out. A healthy rebuild reconstructs a failed disk’s data and writes it to the replacement, using the other disks as the source. If those other disks are themselves faltering, the rebuild can write wrong data — reconstructed from bad reads — over the array, corrupting good information in the process. Worse, some controllers, when you force a rebuild past errors or re-add a disk in the wrong order, will begin overwriting the very parity and data a recovery would need. Because these are writes to the array, they’re often irreversible. That’s why the single most valuable thing you can do with a failed rebuild is nothing further.
If a rebuild has failed or stalled, power the array down and leave it. Don’t initialise, don’t “repair”, don’t swap in more disks hoping one takes, and don’t let the controller start another rebuild on boot. If you can, note the order the disks were in — which bay each came from — because the array can only be reassembled correctly with every member disk in its right place. Then treat the situation as a recovery rather than a repair. The data almost always still exists across the disks; the problem is that the array can no longer safely assemble it itself, and it needs to be done offline instead.
The scenario behind most failed rebuilds is the RAID 5 one: a disk fails, you replace it, the rebuild starts reading all the remaining disks, and a second disk — tired, same age — throws an error partway through, collapsing the rebuild. At that point the array is degraded beyond what it can fix alone, but the data is still recoverable if handled properly. In a lab, every disk is imaged read-only first, so nothing further is written to the originals; the array is then rebuilt virtually from those images, working around the weak disk’s bad sectors, and your files are extracted from the reconstructed volume. It’s the same disciplined approach behind recovering any failed RAID or an inaccessible NAS. Send all the disks labelled in bay order; RAID recovery is from £500 +VAT.
Usually because a second disk is weak and fails under the strain of the rebuild, the controller has lost or muddled the array configuration, an incompatible replacement disk was used, or unreadable sectors are halting the parity calculation. RAID disks tend to age together, so a second disk failing during a rebuild is common. Trying again does not fix any of these and can make things worse.
No. Forcing a rebuild past errors, or repeating a failed one, writes to the array while leaning hard on disks that are already struggling. It can write reconstructed-from-bad-data over good information, and some controllers begin overwriting the parity and data a recovery would need. Because these are writes to the array, the damage is often irreversible.
Power the array down and stop. Do not initialise, repair, swap in more disks, or let the controller start another rebuild on boot. If you can, note which bay each disk came from. Then treat it as a recovery, not a repair, the data almost always still exists across the disks, but the array needs to be reassembled offline rather than by itself.
Usually, yes. The data typically still exists across the member disks; the array simply cannot safely assemble it after a failed rebuild. In a lab, every disk is imaged read-only, the array is rebuilt virtually from those images working around any bad sectors, and the files are extracted. The key is to stop before forcing further rebuilds. Recovery is from 500 pounds plus VAT.
Send all the member disks, labelled with the bay or order they came from, the array can only be reassembled correctly with every disk in its right place. It also helps to know the RAID level (RAID 5, 6, 10 and so on) and the controller or NAS make and model if you have them. A free diagnostic and fixed quote come before any chargeable work.
Don’t force another rebuild. Power the array down and send the disks labelled in order; we rebuild it offline from read-only images — free diagnostic, fixed quote first.