Second Drive Failed During a RAID Rebuild: Stop and Preserve the Set
A second reported drive failure during rebuild means the array has lost another member while every surviving source was under sustained read load. The partial rebuild target, original failed drive, and newly reported failure may each contain different usable regions.
Stop the rebuild. Do not restart with another target, discard either failed member, or force a drive online. Preserve every original and replacement drive exactly as it is now.
Tell an ADR technician what the system reports and what has already been attempted. The failed system does not have to be brought online first.
What this condition means
Rebuild activity reads the full surviving set and can expose weak sectors, timeouts, or intermittent members that normal operation had not revealed.
A controller may label the second member failed after read errors or loss of communication; that label does not establish that the drive is completely unreadable. Meanwhile, the target contains a partial reconstruction that may remain useful.
What to record
- RAID level, controller, enclosure, and member count
- Which drive failed first and which replacement was installed
- Rebuild percentage and exact second-failure messages
- Slot and serial number of every original and replacement member
- Any power cycle, force-online, consistency check, or additional rebuild attempt
What not to do
- Restart the rebuild with another replacement
- Discard the first failed member or partial target
- Force either reported failure online
- Run consistency or filesystem repair
- Write new data to any volume the controller manages to present
How ADR evaluates it
ADR evaluates each member independently and identifies which regions remain readable before another operation writes to the set.
Protected images can combine readable coverage from the original members and partial target. Reconstruction is then tested outside the controller and validated against the filesystem and priority data.
Keeping every member gives ADR the widest evidence set. The server does not need to boot and the controller does not need to recognize the array before imaging and reconstruction can begin.
What happens next
Recoverability depends on whether the combined members still provide enough readable coverage and how much the rebuild changed. Preserving all sources is the most important action now.