← All incidents
Incident dossier · Rank #9

AWS S3 US-EAST-1 Service Disruption

Amazon Web Services 2017-02-28 4h 35m core impact Human errorSoftware

An authorized S3 engineer running an established capacity-removal playbook mistyped one input; the command removed far more servers than intended, taking down the S3 index and placement subsystems in US-EAST-1. Because a very large number of services (and AWS's own status dashboard) depend on S3, the ~4.5-hour disruption cascaded across a large fraction of the internet.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Human error (2017-02-28)Trigger · Human error2017-02-282017-02-28Primary fault at Amazon Web Services — US-EAST-1 region (Northern Virginia)AWS S3US-EAST-1 region (Northern Virginia)US-EAST-1 regionDownstream service degraded by the fault: Amazon S3 (GET/LIST/PUT/DELETE)Amazon S3Downstream service degraded by the fault: EC2 instance launchesEC2 instance launchesDownstream service degraded by the fault: EBS volumesEBS volumesDownstream service degraded by the fault: AWS LambdaAWS Lambda+1 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Amazon Web Services
Data center
US-EAST-1 region (Northern Virginia)
Location
Ashburn (approx.), United States, US-EAST-1
Date
2017-02-28

Impact & scale

Users affected
Thousands of websites and services worldwide dependent on S3 US-EAST-1
Financial
≈ US$150M+ estimated loss to S&P-500 companies (Cyence, via press)
Scope
Hyperscale public-cloud region
Services / systems down
  • Amazon S3 (GET/LIST/PUT/DELETE)
  • EC2 instance launches
  • EBS volumes
  • AWS Lambda
  • AWS Service Health Dashboard (couldn't update — hosted on S3)

Impact data & metrics

Trigger to first partial recovery (index serving GET/LIST/DELETE)~2 h 49 min (09:37 to 12:26 PST)
Trigger to index subsystem full recovery~3 h 41 min (09:37 to 13:18 PST)
Trigger to full S3 restoration (placement/PUT restored)~4 h 17 min (09:37 to 13:54 PST)
Duration status could not be updated on the Service Health Dashboard~2 h (09:37 event start to 11:37 PST); icons 'stuck on green'
S3 subsystems requiring a full restart2 (index + placement)
Time since the index/placement subsystems were last fully restarted in larger regions'many years' (exact interval not disclosed)
Named major external sites/services reported down11+ (Docker Registry Hub, Trello, Travis CI, GitHub, GitLab, Quora, Medium, Signal, Slack, Imgur, Twitch.tv)
AWS regions affected1 (US-EAST-1, Northern Virginia)
Servers intended vs. actually removed / capacity thresholdsNot published — AWS states only that 'a larger set of servers was removed than intended'
Authoritative dollar-loss figureNot disclosed in the primary source; only third-party estimates exist (uncertain)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 9Users affected (0–10) — breadth of the user/customer population impacted. — scored 9/10.Financial 8Financial impact (0–10) — direct + consequential cost. — scored 8/10.Duration 6Outage duration (0–10) — how long service was degraded/down. — scored 6/10.Blast 10Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 10/10.
Magnitude 8.6 = blast 10×0.35 + users 9×0.25 + financial 8×0.20 + duration 6×0.20 (sub-scores 0–10 · weighted composite)

Blast radius extreme: S3 is a foundational dependency for a large share of the public internet; cascade reached EC2/EBS/Lambda and even AWS's own status page.

Sequence of events (SOE)

Phased sequence of events2017-02-28 09:37 PST · TRIGGER — An authorized S3 team member, using an established playbook, executes a command intended to remove a small number of servers from the S3 subsystem used by the S3 billing process.TRIGGER2017-02-28 09:37 PS2017-02-28 09:37 PST · TRIGGER — One input to the command is entered incorrectly; a far larger set of servers is removed than intended.TRIGGER2017-02-28 09:37 PS2017-02-28 ~09:37 PST · DETECTION — The inadvertently removed servers are found to have supported two other S3 subsystems — the index subsystem (metadata/location for all objects) and the placement subsystem (new-storage allocation, dependent on index). The fault is internal and effectively immediate; no external alarm.DETECTION2017-02-28 ~09:37 PS2017-02-28 ~09:38 PST · IMPACT — Index and placement capacity drop below operational minimum; each subsystem now requires a full restart. There was no guardrail to stop removal below the minimum required level.IMPACT2017-02-28 ~09:38 PS2017-02-28 ~09:38 PST · IMPACT — With the subsystems restarting, S3 in US-EAST-1 is unable to service GET, LIST, PUT and DELETE requests.IMPACT2017-02-28 ~09:38 PS2017-02-28 morning PST · CASCADE — S3-dependent AWS services degrade: the S3 console, new EC2 instance launches, EBS operations relying on S3 snapshots, and AWS Lambda are impaired.CASCADE2017-02-28 morning PS2017-02-28 morning PST · CASCADE — External sites go down, including Docker Registry Hub, Trello, Travis CI, GitHub, GitLab, Quora, Medium, Signal, Slack, Imgur and Twitch.tv.CASCADE2017-02-28 morning PS2017-02-28 morning PST · CASCADE — The Service Health Dashboard admin console, itself dependent on S3, cannot update per-service status; status icons are 'stuck on green lights' because the red warning icons were hosted in the downed systems.CASCADE2017-02-28 morning PS2017-02-28 ~09:40 PST · MITIGATION — Recovery begins with a full restart of the index subsystem first, because the placement subsystem depends on it functioning.MITIGATION2017-02-28 ~09:40 PS2017-02-28 ongoing PST · MITIGATION — Metadata-integrity safety checks are run during restart and deliberately not shortcut; at S3's grown scale the restart-plus-validation 'took longer than expected.'MITIGATION2017-02-28 ongoing PS2017-02-28 until 11:37 PST · CASCADE — AWS remains unable to update individual service status on the SHD; it communicates via the @AWSCloud Twitter feed and SHD banner text until the dashboard can be updated (regained around noon PST).CASCADE2017-02-28 until 11:37 PS2017-02-28 12:26 PST · RECOVERY — Index subsystem has activated enough capacity to begin servicing S3 GET, LIST and DELETE requests.RECOVERY2017-02-28 12:26 PS2017-02-28 13:18 PST · RECOVERY — Index subsystem fully recovered; GET, LIST and DELETE APIs functioning normally.RECOVERY2017-02-28 13:18 PS2017-02-28 13:54 PST · RECOVERY — Placement subsystem recovery completes, restoring PUT and new-object storage.RECOVERY2017-02-28 13:54 PS2017-02-28 13:54 PST · RESTORED — All S3 APIs operate normally in US-EAST-1; full service restored roughly 4 hours 17 minutes after the mistyped command.RESTORED2017-02-28 13:54 PS

Root cause

There was no physical ignition source in this incident — no equipment make, model, chemistry, or combustion, and therefore no detection panel, suppression system, evacuation, or emergency-services response. The "ignition" was a human operational error. At 9:37AM PST on 2017-02-28, an authorized S3 team member, following an established debugging playbook, executed a capacity-removal command intended to take a small number of servers out of the S3 subsystem used by the S3 billing process. Per AWS: "one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended." That single mistyped input is the proximate cause. The failure mechanism is a cross-subsystem blast radius the operator did not anticipate. The inadvertently removed servers also supported two other S3 subsystems — the index subsystem, which "manages the metadata and location information of all S3 objects in the region" and is necessary to serve GET, LIST, PUT and DELETE requests, and the placement subsystem, which "manages allocation of new storage and requires the index subsystem to be functioning properly to correctly operate." Dropping their capacity below the operational minimum required each subsystem to be fully restarted, and "while these subsystems were being restarted, S3 was unable to service requests." Recovery was ordered by dependency: index first, then placement. The latent root is procedural/tooling rather than physical, and three-fold. First, a GUARDRAIL gap: the capacity-removal tool accepted an over-broad input with no safeguard against taking any subsystem below its minimum required capacity — AWS's own remediation adding exactly that safeguard is the admission it did not previously exist. Second, an UNTESTED-AT-SCALE recovery procedure: AWS had "not completely restarted the index subsystem or the placement subsystem in our larger regions for many years," and after massive growth the restart plus metadata-integrity validation "took longer than expected" — a cold-start path effectively unrehearsed at current scale. Third, a DEPENDENCY lapse in the status-reporting plane: the Service Health Dashboard admin console itself depended on S3 in the same region, so operators could not update per-service status during the very event it was meant to report.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Nature of the incident: operational error, not a physical fire

This is a cloud-service disruption with no ignition, equipment, chemistry, or combustion — the standard fire-incident lenses (detection panel, suppression, evacuation, de-energisation, emergency-services response) do not apply and are honestly marked N/A. The proximate cause is a single mistyped input to an established capacity-removal playbook at 9:37AM PST on 2017-02-28. AWS states plainly: 'one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended.' The fault was self-inflicted and detected internally and immediately, so there is no alarm latency to analyse — the analysis instead centres on why a scoped command produced a regional outage and why recovery took over four hours.

The blast radius: hidden cross-subsystem coupling

The command targeted a low-privilege pool serving the S3 billing process, but the removed servers also supported two request-critical subsystems. The index subsystem 'manages the metadata and location information of all S3 objects in the region' and is required for every GET, LIST, PUT and DELETE; the placement subsystem allocates new storage and 'requires the index subsystem to be functioning properly to correctly operate.' Over-removal pushed both below their operational minimum with no floor to stop it, forcing full restarts during which 'S3 was unable to service requests.' Scoping by operator intent ('billing') was not scoping by effect.

Why recovery was slow: an unrehearsed cold start at scale

AWS had 'not completely restarted the index subsystem or the placement subsystem in our larger regions for many years.' After massive growth, the cold start plus mandatory metadata-integrity validation 'took longer than expected.' AWS made the correct call to NOT shortcut the safety checks, trading recovery speed for metadata integrity. Recovery followed the dependency chain: index began serving GET/LIST/DELETE at 12:26PM, index fully recovered at 1:18PM, and placement restored PUT/new-object storage at 1:54PM — full restoration roughly 4h17m after the trigger.

The visibility failure: monitoring that depended on the monitored

The most instructive secondary failure was reflexive. The Service Health Dashboard admin console itself depended on S3 in the same region, so 'from the beginning of this event until 11:37AM PST, we were unable to update the individual services' status.' The Register adds that the icons 'were stuck on green lights because the red icons warning of failures were hosted in the downed systems.' AWS fell back to the @AWSCloud Twitter feed and dashboard banner text, and regained dashboard control around noon PST. A status plane must never share a failure domain with the service it reports on.

Systemic lessons and AWS's remediation

AWS committed to concrete, sourced fixes: the capacity-removal tool now removes capacity more slowly and enforces a minimum-capacity safeguard (the admission that no such guardrail existed); the monolithic index subsystem is being partitioned into cells to bound blast radius; and SHD administration was moved to run across multiple regions. The durable lessons generalise beyond AWS: destructive tooling needs hard floors rather than operator care, recovery paths decay silently and must be rehearsed at production scale, and blast-radius partitioning is a design discipline. External casualties (GitHub, Slack, Trello, Docker Registry Hub, Twitch and more) underscore how concentrated a dependency US-EAST-1 S3 had become.

Technical deep-dive

S3 in US-EAST-1 is decomposed into subsystems; two are load-bearing for this incident. The index subsystem is the metadata/location store — it "manages the metadata and location information of all S3 objects in the region" and is necessary to serve all GET, LIST, PUT and DELETE requests. The placement subsystem allocates storage for new objects and "requires the index subsystem to be functioning properly to correctly operate" — a hard directional dependency that dictated recovery order. A third, lower-privilege capacity pool served the S3 billing process; that pool was the intended target of the 9:37AM PST playbook command. The over-removal drove both index and placement below their operational minimum, and there was no floor to stop it. Removing a significant portion of the capacity caused each subsystem to require a full restart, and "while these subsystems were being restarted, S3 was unable to service requests." Because the two subsystems had not been completely restarted "in our larger regions for many years," and S3 had grown massively, the cold-start plus metadata-integrity validation ran far longer than a rehearsed procedure would predict. AWS deliberately did NOT shortcut the safety checks — a correct call trading recovery speed for metadata integrity — which is why the restart and validation "took longer than expected." The cascade radiated across everything treating S3 as a storage primitive: the S3 console, new EC2 instance launches, EBS operations relying on S3 snapshots, and AWS Lambda were impaired. Per The Register, external casualties included Docker Registry Hub, Trello, Travis CI, GitHub, GitLab, Quora, Medium, Signal, Slack, Imgur and Twitch.tv. The most instructive containment failure was reflexive: the AWS Service Health Dashboard admin console itself depended on S3, so "from the beginning of this event until 11:37AM PST, we were unable to update the individual services' status," and per The Register the icons "were stuck on green lights because the red icons warning of failures were hosted in the downed systems." AWS fell back to the @AWSCloud Twitter feed and SHD banner text. Recovery then proceeded along the dependency chain: index began servicing GET/LIST/DELETE at 12:26PM PST, index fully recovered at 1:18PM PST, and placement (restoring PUT/new-object storage) completed at 1:54PM PST, at which point S3 operated normally. The record is explicit that this was a self-inflicted fault detected internally and immediately — no alarm latency, no external responders.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home