Call us — 01483 901310
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →

Data Recovery Case File · NAS & Network Storage · Scale Changes the Odds

Large Members Make Single-Parity Arrays Fragile

This enquiry came from a research institution and the numbers in it are the case. A single-parity array "consisting of 12 drives, with storage capability of 7.3 TB each" — roughly seventy-eight terabytes of usable capacity — holding sequencing data. Single parity across members of that size is a configuration with a known statistical weakness, and it is not about any drive being unreliable. It is about how long a rebuild takes and what happens during it.

MediaTwelve enterprise-class drives of approximately 7.3TB each in a single-parity array — approximately 78TB usable; scientific dataset held
Reported situationResearch server operating a single-parity array of twelve large members · substantial scientific dataset held · array in a failed or degraded state · recovery of dataset required
Fault classSingle-parity array with very large members — rebuild exposure proportional to member size; offline reconstruction indicated
Equipment usedNo rebuild attempted in the array · all members imaged individually write-blocked (Atola TaskForce 2) · failure sequence established from member metadata · volume assembled offline from images · output restored to separate hardware

The decode: why size, not reliability, is the problem

What single parity provides: tolerance of one member failing. The array continues in a degraded state, and when a replacement is fitted it reconstructs the missing member by reading every sector of every surviving drive and recalculating.

Why that operation is the weak point: a rebuild across eleven surviving members of 7.3TB each means reading roughly eighty terabytes, continuously, at maximum load, with no redundancy remaining. On drives of that size the operation takes many hours and frequently days. Throughout all of it, a second failure is unrecoverable.

Why the risk scales with capacity rather than count: drives have not become proportionally more reliable per byte as they have grown. A rebuild's exposure is a function of how much must be read without error — and reading eighty terabytes presents far more opportunities for an unrecoverable read than reading eight. Single parity was designed when members were a fraction of this size, and the arithmetic has moved underneath it.

What that means practically for a configuration like this: the array is fully protected against one failure and effectively unprotected during the recovery from one. The period of greatest vulnerability is the period immediately after the first thing goes wrong, which is precisely when people are relying on the redundancy to save them.

Why a rebuild must not be attempted now: the same reasons, amplified. Every surviving member is the same age, from the same batch, with the same workload as whichever failed. Asking all eleven to perform their heaviest ever read simultaneously, unprotected, is how a partial loss becomes total.

What is done instead: every member is imaged individually, once, under controlled conditions — including any that have failed. The failure sequence is established from the metadata each drive carries, so stale members can be excluded rather than mixed in, and the volume is assembled offline from the copies. Same information, no load on any survivor beyond a single controlled read, and the assembly can be attempted repeatedly at no cost.

What the scale adds practically: imaging seventy-eight terabytes is a substantial undertaking in time and destination capacity, and it should be planned rather than started. That is a logistics conversation as much as a technical one.

What to change afterwards: at these member sizes, double parity is the minimum sensible configuration — because it keeps redundancy during the rebuild, which is the moment that actually matters.

On the bench

No rebuild was attempted in the array — reconstruction requiring every sector of every surviving member to be read at maximum load with no redundancy, which across eleven drives of this size runs for days and is precisely when a second failure becomes unrecoverable. All members were imaged individually write-blocked on the Atola TaskForce 2, the failure sequence established from member metadata so stale members could be excluded, and the volume assembled offline from the images with output restored to separate hardware.

The outcome

Every member imaged individually, the failure sequence established from metadata and the volume assembled offline onto separate hardware. Free assessment, one fixed written figure including VAT; where a drive has to be opened, 50% of parts and labour is payable upfront with the balance only on success — otherwise no recovery, no fee. The decode: single parity tolerates one failure and rebuilds by reading every sector of every survivor — which across members this large takes days at full load with no protection remaining. The vulnerable period is the recovery from the first failure, not the failure itself.

Large single-parity array in a degraded state

Don't start a rebuild. Single parity tolerates one member failing, and it recovers by reading every sector of every surviving drive and recalculating — which across members of this size means reading tens of terabytes continuously, for days, at maximum load, with no redundancy left. Your remaining drives are the same age and batch with the same workload as the one that went, and asking all of them to perform their heaviest ever read unprotected is how a partial loss becomes total. The exposure scales with capacity rather than drive count, because drives haven't become proportionally more reliable per byte. Image every member individually instead and assemble offline. Afterwards, double parity.

Large array degraded and facing a rebuild?
Don't rebuild in place — call Guildford Data Recovery on 01483 901310; every member imaged individually, failure sequence established from metadata, volume assembled offline onto separate hardware.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.