Data Recovery Case File · NAS & Network Storage · Assemble Offline Instead
The Failed Disk Is the Key, and the Array Should Not Turn It
This enquiry came from an IT technician and proposes a specific, sensible plan. An array where "one more has failed than it can recover from. Is it possible for you to recover one of the failed disks so we can replace it and allow the array to recover itself?" The reasoning is sound — one working member short, so restore one member. The instinct about where the answer lies is exactly right, and the part to change is who does the work, because letting the array rebuild is repeating the operation that lost the second disk.
| Media | Redundant array with more failed members than its parity level tolerates — all members retained; reconstruction sought |
| Reported situation | Array in a failed state · failures exceeding the configuration's tolerance by one member · all disks retained · proposal to restore one failed member and allow the array to rebuild · volume required |
| Fault class | Multiple member failure beyond parity — offline reconstruction indicated; in-array rebuild contraindicated by load and by staleness |
| Equipment used | No member returned to the array and no rebuild attempted · all members including both failed disks imaged individually write-blocked (Atola TaskForce 2) · failure sequence established from member metadata · volume assembled offline from images · output restored to separate hardware |
The decode: the instinct is right, and two things make the method wrong
Why the reasoning is sound: an array one member short of viability needs one more member's worth of data, and the failed disk holds it. Recovering that disk does supply exactly what is missing. As an analysis of where the information lives, it is correct.
The first problem — the rebuild is the load that killed the last disk. Returning a restored member and letting the array reconstruct means reading every sector of every surviving member, continuously, for hours or days. That is the heaviest work the array ever does, and it happens with no redundancy left. The remaining disks are the same age, from the same batch, with the same workload as the ones that already failed. Asking them to do the hardest possible job at the least protected moment is how a partial loss becomes a total one — and it is precisely how the second failure usually occurs in the first place.
The second — a restored member is stale. A disk that dropped out holds the array's state as of the moment it failed, not as of now. If the array continued running afterwards, that disk's contents are out of date relative to the others. Returning it means the controller treating old data as current and reconstructing from it, which produces a volume that mounts, looks healthy, and contains silently corrupted files. That is a worse outcome than a volume that will not mount, because nothing announces it.
What is done instead, using the same insight: every member is imaged individually, including both failed disks, so each is read once under controlled conditions while it still cooperates. The failure sequence is then established from the metadata each disk carries — which records the array's state and a generation number — so stale members can be identified and excluded rather than mixed in. The volume is assembled offline from the images.
Why that is strictly better: it uses the same information the rebuild would have used, applies no load to any surviving disk beyond a single controlled read, allows the assembly to be attempted repeatedly at no cost, and produces output on separate hardware rather than leaving the result on the disks that failed.
The practical note for afterwards: the restored volume should go to new hardware. Disks that reached this point together will reach the next point together too.
On the bench
No member was returned to the array and no rebuild was attempted — a reconstruction reads every sector of every survivor at maximum load with no redundancy remaining, which is the operation that produces second failures, and a restored member holds the array's state from the moment it dropped out rather than the present. All members including both failed disks were imaged individually write-blocked on the Atola TaskForce 2, the failure sequence established from member metadata so stale disks could be excluded rather than mixed in, and the volume assembled offline from the images with output restored to separate hardware.
The outcome
Every member imaged individually, the failure sequence established from metadata and the volume assembled offline onto separate hardware. Free assessment, one fixed written figure including VAT; where a drive has to be opened, 50% of parts and labour is payable upfront with the balance only on success — otherwise no recovery, no fee. The decode: the instinct is right — the failed disk holds what the array is missing. Two things make the in-array method wrong. A rebuild reads every survivor end to end at full load with no redundancy left, which is how second failures happen. And a restored member is stale, so reconstructing from it yields a volume that mounts with silently corrupted contents.
Array beyond its parity level with disks retained
Keep every disk including both failed ones, and don't return a repaired member to the array. Your reasoning about where the missing information lives is correct — the failed disk does hold it. The problem is the method. A rebuild reads every sector of every surviving member continuously, at maximum load, with no redundancy left to absorb anything, and your remaining disks are the same age and batch with the same workload as the ones that already went. That operation is how second failures happen rather than a way to fix them. There's also staleness: a disk that dropped out holds the array's state from that moment, so reconstructing from it can produce a volume that mounts and contains silently corrupted files.
Don't rebuild in place — call Guildford Data Recovery on 01483 901310; every member imaged individually, failure sequence established from metadata, volume assembled offline onto separate hardware.
Request a quote online →
Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.