← All incidents
Incident dossier · Rank #22

AWS EBS Re-Mirroring Storm and Stuck Volumes in US-East-1

Amazon Web Services 2011-04-21 90h 43m core impact Networkconfig-errorSoftwarestorage

A routine network-capacity upgrade in one US-East-1 Availability Zone was executed incorrectly, shifting EBS traffic onto a low-capacity redundant network and isolating a large block of storage nodes. When connectivity returned, the nodes launched a self-reinforcing "re-mirroring storm" that exhausted spare capacity, left ~13% of the zone's volumes "stuck", and starved the region-wide EBS control plane — degrading EC2 and RDS for nearly four days and permanently losing 0.07% of the affected zone's volumes.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2011-04-21)Trigger · Network2011-04-212011-04-21Primary fault at Amazon Web Services — US-East-1 (Northern Virginia)AWSUS-East-1 (Northern Virginia)US-East-1Downstream service degraded by the fault: Amazon EBS (Elastic Block Store)Amazon EBSDownstream service degraded by the fault: Amazon EC2 (EBS-backed instance launch/attach)Amazon EC2Downstream service degraded by the fault: Amazon RDS (single-AZ and some multi-AZ)Amazon RDS

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Amazon Web Services
Data center
US-East-1 (Northern Virginia)
Location
Ashburn area, Northern Virginia, USA, single affected Availability Zone (of the US East Region)
Date
2011-04-21

Impact & scale

Users affected
Not disclosed as a user count; blast radius covered one AZ directly plus region-wide EBS/EC2/RDS control-plane degradation. Consumer sites knocked offline included Reddit, Quora, Foursquare, Hootsuite and Heroku (also Engine Yard), per CNN Money and GovTech.
Financial
Not disclosed by AWS. Direct remedy was a 10-day service credit equal to 100% of affected customers' EBS, EC2 and RDS usage in the affected AZ; aggregate dollar figure never published.
Scope
Major multi-day regional service disruption (single AZ hard-down, region-wide control-plane impact)
Services / systems down
  • Amazon EBS (Elastic Block Store)
  • Amazon EC2 (EBS-backed instance launch/attach)
  • Amazon RDS (single-AZ and some multi-AZ)

Impact data & metrics

Total disruption duration~90 hours (approx; 12:47 AM PDT Apr 21 to effective end over the weekend of Apr 24; AWS gives no precise closing time)
Peak stuck EBS volumes (affected AZ)~13% initial, +5% in second wave (~18% cumulative touched)
Volumes not yet restored by Apr 22 12:30 PDT2.2%
Volumes not yet restored by Apr 24 12:30 PDT1.04%
Permanent EBS data loss (affected AZ)0.07% of volumes unrecoverable
RDS single-AZ instances with stuck I/O (peak)45% at peak; 0.4% suffered unrecoverable storage
RDS single-AZ recovery pace41.0% still stuck after 24h, 23.5% after 36h, 14.6% after 48h; the rest recovered over the weekend
RDS multi-AZ failover failures2.5% failed to auto-failover due to a previously un-encountered bug
Stuck-volume spillover into healthy AZs (peak)less than 0.07%
Customer remedy10-day service credit = 100% of affected EBS, EC2 and RDS usage in the affected AZ

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 9Outage duration (0–10) — how long service was degraded/down. — scored 9/10.Blast 7Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 7/10.
Magnitude 7.5 = blast 7×0.35 + users 8×0.25 + financial 6×0.20 + duration 9×0.20 (sub-scores 0–10 · weighted composite)

Duration score high: ~90 hours (approx) from trigger (12:47 AM PDT Apr 21) to effective end over the weekend of Apr 24. AWS's own post-mortem states RDS databases 'recovered throughout the weekend' without a precise closing time; third-party summaries cite ~7:30 PM PDT Apr 24. Blast radius: one AZ hard-down with region-wide EBS control-plane thread starvation touching all AZs; peak stuck-volume spillover to healthy zones stayed <0.07%. Financial mid: real remedy was a 10-day 100% EBS/EC2/RDS credit in the affected AZ but AWS never published a dollar total.

Sequence of events (SOE)

Phased sequence of events2011-04-21 00:47 PDT · TRIGGER — During a planned network-capacity upgrade in one US-East-1 AZ, engineers shifted traffic off a primary EBS router; instead of moving it to another router in the primary high-capacity network, the change routed it onto the lower-capacity redundant EBS network, which could not carry the load.TRIGGER2011-04-21 00:47 PD2011-04-21 00:47 PDT · TRIGGER — Because the redundant network also became overwhelmed/isolated, affected EBS nodes lost BOTH their primary and secondary network paths at once, cutting them off from their replica peers.TRIGGER2011-04-21 00:47 PD2011-04-21 ~00:50 PDT · CASCADE — The incorrect traffic shift was rolled back; reconnected nodes immediately began searching the cluster for free space to re-establish lost data replicas, all at once. This concurrent scramble became a self-reinforcing 're-mirroring storm'.CASCADE2011-04-21 ~00:50 PD2011-04-21 ~01:00 PDT · CASCADE — Free capacity in the affected cluster was rapidly exhausted; nodes that could not find space entered a retry loop. About 13% of the volumes in the affected AZ became 'stuck' searching for storage.CASCADE2011-04-21 ~01:00 PD2011-04-21 ~01:15 PDT · CASCADE — The flood of Create Volume / re-mirror API calls caused thread starvation in the EBS control plane. Because the control plane spans the whole region, elevated error rates and latencies appeared across ALL Availability Zones, not just the failed one.CASCADE2011-04-21 ~01:15 PD2011-04-21 ~02:00 PDT (just before 5 a.m. ET) · IMPACT — Consumer-facing sites hosted on EC2 in the region went down or degraded, including Reddit, Quora, Foursquare, Hootsuite and Heroku (also Engine Yard); the outage was public and widely reported. Press reports place the onset just before 5 a.m. ET, i.e. roughly 2 a.m. PDT.IMPACT2011-04-21 ~02:00 PD2011-04-21 02:40 PDT · MITIGATION — AWS disabled all Create Volume API requests in the affected AZ to relieve the control-plane thread starvation.MITIGATION2011-04-21 02:40 PD2011-04-21 02:50 PDT · MITIGATION — With Create Volume calls blocked, EBS control-plane latencies and error rates for all other EBS-related APIs in the OTHER (healthy) Availability Zones recovered to normal.MITIGATION2011-04-21 02:50 PD2011-04-21 05:30-08:20 PDT · CASCADE — A second wave hit: a low-probability race condition caused EBS nodes to fail when concurrently closing large numbers of replication requests, and failed re-mirror retries generated heavy control-plane negotiation, pushing an ADDITIONAL ~5% of the AZ's volumes into the stuck state.CASCADE2011-04-21 05:30-08:20 PD2011-04-21 08:20 PDT · MITIGATION — AWS began disabling all communication between the degraded EBS cluster and the EBS control plane, quarantining the failure so it could no longer poison other zones.MITIGATION2011-04-21 08:20 PD2011-04-21 11:30 PDT · RECOVERY — A change to the EBS control plane fixed the elevated errors/latencies when launching EBS-backed EC2 instances in the healthy zones; new-instance error rates declined rapidly and returned to near-normal by Noon PDT.RECOVERY2011-04-21 11:30 PD2011-04-21 12:04 PDT · RECOVERY — The event was contained to the single affected AZ: stabilization complete, ~13% of the AZ's volumes still stuck, and EBS APIs disabled in the affected zone while recovery proceeded.RECOVERY2011-04-21 12:04 PD2011-04-22 02:00 PDT · RECOVERY — Installation of new physical storage capacity began. Recovery was slow because the EBS cluster will not reuse a failed node until every replica on it is successfully re-mirrored, and capacity had to be physically relocated from across the US East Region.RECOVERY2011-04-22 02:00 PD2011-04-22 12:30 PDT · RECOVERY — All but about 2.2% of the volumes in the affected AZ had been restored to normal operation.RECOVERY2011-04-22 12:30 PD2011-04-23 15:35 PDT · RECOVERY — EBS control-plane API access was re-enabled (finished enabling access) for the affected AZ.RECOVERY2011-04-23 15:35 PD2011-04-23 18:15 PDT · RECOVERY — Full EBS API access to EBS resources was restored to the affected Availability Zone.RECOVERY2011-04-23 18:15 PD2011-04-24 12:30 PDT · RECOVERY — All but 1.04% of the affected volumes had been recovered, largely by restoring from snapshots.RECOVERY2011-04-24 12:30 PD2011-04-24 15:00 PDT · RECOVERY — Manual recovery began on the remaining volumes that had not been snapshotted / were tied to hardware failures.RECOVERY2011-04-24 15:00 PD2011-04-24 (evening, over the weekend) · RESTORED — Remaining RDS databases were brought back online over the weekend, effectively ending the multi-day disruption; a final 0.07% of the affected AZ's EBS volumes could not be recovered. AWS's post-mortem describes RDS recovery as completing 'throughout the weekend' without a precise time; third-party summaries cite ~7:30 PM PDT Apr 24.RESTORED2011-04-24 (even

Root cause

Immediate mechanism: At 12:47 AM PDT on 21 April 2011, AWS engineers ran a planned capacity upgrade on the primary EBS network in one US-East-1 Availability Zone. The procedure required shifting traffic off one router in the primary (high-capacity) EBS network onto another router in that same primary network. Instead, the change was executed so that traffic was routed onto the LOWER-CAPACITY redundant EBS network. That secondary network was never sized to carry primary load, so it too saturated/failed — and because the shift removed the primary path at the same moment, a large group of EBS storage nodes lost BOTH their primary and secondary network connectivity simultaneously. EBS nodes depend on continuous peer-to-peer connectivity to keep each volume's two replicas in sync; losing all connectivity made them believe their replicas had failed. The specific failure that turned a network blip into a four-day outage was the EBS re-mirroring design. When connectivity was restored, every isolated node concurrently and aggressively began scanning the cluster for free space to rebuild a fresh second replica. Because so many volumes re-mirrored at once, the cluster's free capacity was exhausted almost instantly, and nodes that could not find space fell into a tight retry loop — the "re-mirroring storm". Roughly 13% of the AZ's volumes ended up "stuck", unable to serve I/O while hunting for storage. The storm then attacked a second, shared component: the region-wide EBS control plane. The torrent of Create Volume and re-mirror negotiation calls caused thread starvation in the control plane, and because that control plane is a single regional service, its degradation leaked EC2 launch/attach errors and elevated latencies into ALL Availability Zones — converting a single-AZ hardware event into a regional one. A latent, low-probability race condition (nodes failing while concurrently closing large batches of replication requests) added a second wave of ~5% more stuck volumes. Latent / organisational root causes: (1) A change-management gap — a manual, error-prone network migration step with no automated guardrail to detect or block routing production EBS traffic onto a network known to lack the required capacity; the safeguard that should have caught it (capacity-aware validation or a peer-review/dry-run of the traffic-shift target) was absent or bypassed. (2) A design flaw in re-mirroring: it had no back-pressure, jitter, or admission control, so a correlated failure produced a stampede that consumed the very spare capacity meant to absorb failures — a classic positive-feedback loop. (3) A blast-radius / isolation flaw: the EBS control plane was a shared, region-scoped dependency with no per-AZ fault isolation, so one AZ's storm starved the whole region. (4) A hidden defect in the multi-AZ RDS failover path (a "previously un-encountered bug") meant 2.5% of the very databases customers had paid extra to make resilient did not fail over automatically.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Canonical control-plane cascade

a localized data-plane network misconfiguration in one AZ propagated region-wide because the EBS control plane was a single shared regional service subject to thread starvation.

Positive-feedback re-mirroring loop

reconnected nodes all re-mirrored at once, exhausting the free space that re-mirroring itself needed, so the more the system tried to heal the worse it got.

Two stacked amplifiers

(a) the re-mirroring stampede (~13%) and (b) a low-probability race condition on concurrent replication-request closure (~5%).

Blast-radius containment held

Blast-radius containment worked once AWS quarantined the degraded cluster from the control plane (08:20 PDT) and blocked Create Volume (02:40 PDT), restoring healthy AZs while the affected AZ was worked through the weekend.

Data durability held

only 0.07% of the affected AZ's volumes were permanently lost, and most recovery came from customer/EBS snapshots — reinforcing the snapshot discipline lesson.

RDS single-AZ dependency exposed

The RDS impact showed that a dependency on EBS turns an EBS event into a database event, and that the multi-AZ failover safeguard itself had an untested failure mode (2.5% did not fail over).

Technical deep-dive

EBS stores each volume as two replicas on separate storage nodes kept in sync over a peer-to-peer network. Normal operation uses a high-capacity primary network with a lower-capacity redundant network intended only for inter-node control/backup traffic. At 12:47 AM PDT on Apr 21, a capacity-upgrade step meant to move traffic between two routers on the primary network instead routed a large volume of traffic onto the low-capacity redundant network. That network saturated, and because the primary path had been withdrawn simultaneously, a large set of nodes lost all connectivity to their replica peers. Each node, unable to reach its replica, concluded the replica had failed and — upon reconnection — tried to establish a new replica by scanning the cluster for free space. With a large fraction of the cluster doing this at once, free capacity was consumed almost immediately; nodes that could not find space entered a continuous retry loop, leaving ~13% of the AZ's volumes 'stuck' (unable to serve I/O). The re-mirror negotiation and a flood of Create Volume calls then exhausted worker threads in the region-wide EBS control plane, so elevated latency/errors appeared across all AZs. AWS disabled Create Volume in the affected AZ (02:40), which recovered other AZs' EBS APIs by 02:50; but a low-probability race condition — nodes failing while concurrently closing large numbers of replication requests — drove a second wave (~5% more stuck) between 05:30 and 08:20, prompting AWS to sever the degraded cluster from the control plane (08:20). A control-plane fix at 11:30 restored EBS-backed EC2 launches to near-normal by noon; the event was contained to the single AZ by 12:04 PM. Recovery was slow because a failed node cannot be reused until all its replicas are re-mirrored and physical capacity had to be relocated across the region: 2.2% still down by Apr 22 12:30, 1.04% by Apr 24 12:30, with 0.07% permanently lost. RDS, which is layered on EBS, saw a peak 45% of single-AZ instances with stuck I/O (0.4% unrecoverable) and 2.5% of multi-AZ instances failing to auto-failover due to a previously un-encountered bug; RDS recovered over the following weekend.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-08.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home