RAID Degraded and Unstable: What to Do Before Rebuilding
A degraded array has already lost redundancy. If it is also slowing down, dropping members, producing read errors, or making the volume disappear, the surviving set may not be healthy enough to support a rebuild.
Reduce nonessential writes and do not begin a rebuild until the failed member, the remaining members, and the sequence of events are understood.
What this condition means
Degraded describes the RAID membership state. Unstable describes what the remaining storage is doing. The combination can point to unreadable sectors, timeouts, a controller or path problem, inconsistent metadata, or another member beginning to fail.
A rebuild places sustained read load on every required source and writes a new member state. If the sources are not trustworthy, restoring redundancy can turn a recoverable outage into a failed rebuild.
What to record
- RAID level, controller, enclosure, and member count
- Slot, serial number, and displayed status of every drive
- The first alert and any later changes in member status
- Read errors, timeouts, slowdowns, or abnormal sounds
- Every replacement, reboot, rebuild, import, or repair already attempted
What not to do
- Start a rebuild simply because the controller recommends one
- Remove or reseat several members together
- Force a reported member online
- Run consistency or filesystem repair
- Continue heavy production writes while the array is changing state
How ADR evaluates it
ADR separates member failure from controller, enclosure, and metadata problems, then determines whether the active sources are readable enough for a controlled recovery plan.
Stable cases may be evaluated at the client location. If a member is unstable, it is protected before reconstruction proceeds away from the original controller.
ADR can review the failure sequence, connect to stable members independently, and protect any unstable member before reconstruction. The degraded array does not need to remain online or accept another controller write for that work to begin.
What happens next
A degraded array is often recoverable and may still be serving data. The unstable behavior is the reason to preserve the current state before asking the remaining members to rebuild it.