← All incidents
Incident dossier · Rank #30

AWS US-EAST-1 Cooling/Thermal Event and use1-az4 Power Loss

Amazon Web Services 2026-05-07 20h 25m core impact CoolingPower

On 7 May 2026 cooling systems failed in AWS's use1-az4 availability zone in US-EAST-1 (Northern Virginia). As data-hall temperatures crossed operating thresholds, servers automatically shut down to protect hardware, cutting power to the EC2 instances and EBS volumes hosted on that gear. The compute and storage loss cascaded to dependent services — IoT Core, ELB, NAT Gateway and Redshift — with elevated error rates and slower provisioning. AWS shifted traffic off the impaired zone and advised customers to relaunch or redistribute workloads in other US-EAST-1 availability zones. AWS first flagged the problem at 5:25 PM PDT on 7 May and restored cooling to pre-incident capacity at 1:50 PM PDT on 8 May, roughly 20 hours 25 minutes later. Downstream, Coinbase reported a multi-hour outage in which users could not trade or move funds. No formal AWS post-incident report was published in the available sources.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Cooling (2026-05-07)Trigger · Cooling2026-05-072026-05-07Primary fault at Amazon Web Services — US-EAST-1 (use1-az4), Northern VirginiaAWSUS-EAST-1 (use1-az4), Northern VirginiaUS-EAST-1Downstream service degraded by the fault: EC2 (direct)EC2Downstream service degraded by the fault: EBS (direct)EBSDownstream service degraded by the fault: IoT Core (cascaded)IoT CoreDownstream service degraded by the fault: ELB (cascaded)ELB+2 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Amazon Web Services
Data center
US-EAST-1 (use1-az4), Northern Virginia
Location
Ashburn, United States, use1-az4
Date
2026-05-07

Impact & scale

Users affected
Not publicly quantified. Directly impaired EC2/EBS customers in use1-az4 plus downstream users of dependent services; named downstream victim Coinbase suffered a multi-hour trading and funds-movement outage, but no affected-user count was disclosed by any source.
Financial
Not disclosed. No source states a dollar loss for AWS or for downstream customers; Coinbase impact was reported only qualitatively.
Scope
Major single-AZ regional impairment with multi-service cascade and named downstream customer outage
Services / systems down
  • EC2 (direct)
  • EBS (direct)
  • IoT Core (cascaded)
  • ELB (cascaded)
  • NAT Gateway (cascaded)
  • Redshift (cascaded)

Impact data & metrics

Total incident duration (onset to resolved)~28 hours (reported)
Cooling-restoration duration (onset to pre-incident capacity)~21.5 hours (4:20 PM PDT May 7 -> 1:50 PM PDT May 8)
Threshold-exceedance / trigger time4:20 PM PDT May 7 (23:20 UTC; Axis separately cites 23:50 UTC - internal inconsistency)
Coinbase customer downtime~7 hours (reported)
Blast radius1 data hall / 1 Availability Zone (use1-az4), described as a very heavily used AWS zone
Cooling units failed'Multiple' / 'multiple chiller units' near-simultaneously (exact count not disclosed)
Official AWS post-event summary published0 (none for May 2026; latest PES entry is the Oct 19, 2025 DynamoDB disruption)
Temperature detail disclosed (deg / setpoint / threshold value)Not disclosed - only 'exceeded safe operating thresholds', no numeric value
Chiller/CRAH make, model, age, refrigerantNot disclosed in any source
Recency ranking of the outageReported as 4th significant us-east-1 outage since October 2025 (single-source, unverified)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 7Users affected (0–10) — breadth of the user/customer population impacted. — scored 7/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 7Outage duration (0–10) — how long service was degraded/down. — scored 7/10.Blast 6Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 6/10.
Magnitude 6.5 = blast 6×0.35 + users 7×0.25 + financial 6×0.20 + duration 7×0.20 (sub-scores 0–10 · weighted composite)

Scores reflect a long (~20h) single-AZ event whose direct footprint (EC2/EBS in one zone) fanned out to four cascaded AWS services and at least one high-profile downstream customer (Coinbase). User and financial scores are estimates constrained by disclosure: no affected-user count or dollar loss was published, so financial (6) and users (7) are inferred from the scale of the region and named victims rather than measured. Duration (7) is anchored to the two firmly stamped endpoints. Blast radius (6) is moderated by the impairment being confined to one availability zone, not the whole region.

Sequence of events (SOE)

Phased sequence of events2026-05-07 ~16:20 PDT (23:20 UTC; Axis separately cites 23:50 UTC) · TRIGGER — Multiple cooling/chiller units in a single data hall of Availability Zone use1-az4 (Ashburn) reportedly fail near-simultaneously; hall cooling capacity drops sharply.TRIGGER2026-05-07 ~16:20 PD2026-05-07 ~16:20 PDT · DETECTION — Rack- and hall-level temperature sensors register temperatures exceeding safe operating thresholds; detection was thermal/environmental, NOT smoke/fire (no VESDA, no fire alarm reported).DETECTION2026-05-07 ~16:20 PD2026-05-07 ~16:20 PDT · MITIGATION — Servers execute their built-in automatic over-temperature protective shutdown to protect silicon; no fire-suppression system involved.MITIGATION2026-05-07 ~16:20 PD2026-05-07 ~16:20 PDT · IMPACT — The protective shutdown de-energises the affected racks; EC2 instances and EBS volumes lose power, some physically damaged.IMPACT2026-05-07 ~16:20 PD2026-05-07 ~17:06 PDT · MITIGATION — AWS reportedly shifts customer traffic away from use1-az4 to contain the blast radius to the single affected zone.MITIGATION2026-05-07 ~17:06 PD2026-05-07 17:25 PDT · DETECTION — AWS Health Dashboard reportedly posts first status message: 'EC2 instances and EBS volumes hosted on impacted hardware are affected by the loss of power during the thermal event.'DETECTION2026-05-07 17:25 PD2026-05-07 evening PDT · IMPACT — Downstream customers begin outages; Coinbase and other services on affected racks reportedly go dark.IMPACT2026-05-07 evening PD2026-05-07 18:47 PDT · CASCADE — AWS Health Dashboard reportedly warns dependent services will be hit: other AWS services depending on affected EC2/EBS in the AZ 'may also experience impairments.'CASCADE2026-05-07 18:47 PD2026-05-07 evening PDT · CASCADE — Software-orchestration recovery proves impossible for physically damaged hardware: 'Software orchestration cannot automatically reroute around physical hardware damage.'CASCADE2026-05-07 evening PD2026-05-07 22:11 PDT · RECOVERY — AWS reports staged cooling-restoration underway ('incremental progress to restore cooling systems'); recovery deliberately paced to avoid damaging thermally-stressed equipment.RECOVERY2026-05-07 22:11 PD2026-05-08 06:51 PDT · DETECTION — AWS reportedly posts root-cause status message confirming the thermal->shutdown->power-loss chain ('servers automatically shut down ... to protect the hardware').DETECTION2026-05-08 06:51 PD2026-05-08 (during recovery) · RECOVERY — AWS advises customers to restore from EBS snapshots or launch resources in unaffected zones, acknowledging some hardware will not come back in place.RECOVERY2026-05-08 (duri2026-05-08 13:50 PDT · RESTORED — Cooling reportedly restored to pre-incident capacity in the affected hall.RESTORED2026-05-08 13:50 PD2026-05-08 (Coinbase) · RESTORED — Coinbase service reportedly restored after roughly 7 hours of downtime.RESTORED2026-05-08 (Coin2026-05-08 ~20:04 PDT · RESTORED — Incident marked resolved (~28 hours after onset); however some EC2 instances/EBS volumes reportedly remained impaired due to physical damage.RESTORED2026-05-08 ~20:04 PD2026-08-02 · RESTORED — VERIFIED: AWS Post-Event Summaries archive checked; most recent entry is the Oct 19 2025 DynamoDB disruption. No formal post-event summary for this May 2026 thermal event was ever published.RESTORED2026-08-02

Root cause

SPECIFIC PRECIPITATING FAILURE (a cooling/thermal-management failure, NOT a fire). No source reports fire, smoke, flame, fire-alarm activation, suppression discharge, evacuation, or fire-department response. The precipitating failure was the near-simultaneous loss of multiple cooling/chiller units in a SINGLE data hall of Availability Zone use1-az4 (Ashburn / Northern Virginia, us-east-1) at approximately 4:20 PM PDT (23:20 UTC) on 2026-05-07. This attribution rests on secondary trade reporting (SingleStore: "Multiple cooling units in availability zone use1-az4 failed"; Axis Intelligence: "multiple chiller units failed simultaneously in a single data hall") that COULD NOT be re-fetched this review session (search budget exhausted, no URLs supplied), so it is UNVERIFIED here.\n\nFAILURE MECHANISM (physical chain, as reported): loss of cooling capacity -> rapid rise of rack-inlet temperature -> temperatures crossed safe operating thresholds -> servers executed their built-in automatic over-temperature protective shutdown to protect silicon -> that protective shutdown removed power from the affected racks -> EC2 instances and EBS volumes lost power, with some hardware physically damaged. This chain is attributed to AWS Health Dashboard / AWS Builder Center wording ("servers automatically shut down when the temperatures exceeded the operating thresholds in order to protect the hardware") reproduced by third parties; the primary AWS text was NOT independently accessible this session.\n\nCRITICAL UNKNOWNS — the record stops short of component-level root cause. No source discloses the chiller/CRAH/CDU make, model, age, refrigerant chemistry, or the specific mechanical/electrical fault that caused the units to fail together. "Simultaneously" implies a COMMON-CAUSE failure (shared chilled-water loop, shared electrical feed to the cooling plant, or a shared BMS/controls fault) but none is confirmed. VERIFIED FACT: AWS published no formal Post-Event Summary for this incident — the official AWS Post-Event Summaries archive's most recent entry, accessed 2026-08-02, is the October 19, 2025 DynamoDB disruption; there is NO May 2026 cooling/thermal summary. The authoritative record therefore stops at (unverified reproductions of) AWS Health Dashboard status messages.\n\nLATENT ROOT (design/redundancy, inferred not disclosed). The evidenced forensics point to a redundancy-and-concentration latent root, not a maintenance record: cooling redundancy within the hall was insufficient to survive the loss of multiple units at once, and use1-az4 is described as a very heavily used zone, concentrating load behind a single physical hall. MAINTENANCE/INSPECTION/TESTING: NOT DISCLOSED (state as "not disclosed," never as "no lapse occurred"). Because AWS published no PES, no maintenance record, work-order, preventive-maintenance status, or BMS-configuration history was released. Whether the simultaneous failure stemmed from deferred maintenance, a failed PM action, a controls misconfiguration, a common utility/coolant fault, or unmaintainable load growth is genuinely unknown.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What actually failed - and what did not

This was a cooling/thermal-management failure in a single data hall of use1-az4, NOT a fire. No ignition, smoke detection, suppression discharge, evacuation, or fire-department response is reported anywhere. Multiple cooling/chiller units reportedly failed near-simultaneously; the mechanical component, make, model, age, refrigerant, and specific fault were never disclosed. 'Simultaneously' implies a common-cause failure (shared coolant loop, shared power feed, or a controls fault) but none is confirmed. All component-level detail remains a disclosed gap.

Why a safety mechanism manufactured the outage

When cooling collapsed, rack-inlet temperatures crossed safe operating thresholds and servers executed their built-in over-temperature protective shutdown to protect silicon. That shutdown, working exactly as designed, de-energised the affected racks - and the power loss physically damaged some EC2 instances and EBS volumes. The customer-visible outage was therefore produced by protection succeeding, not failing. This is the pivotal, counter-intuitive forensic point and is the correct frame for the whole incident.

The physical-damage recovery tail

Recovery was an engineering cooling-restoration effort, deliberately staged because a rushed re-energisation could thermally cycle and damage already-stressed hardware. Cooling was reportedly restored to pre-incident capacity ~21.5 hours after onset, with the incident resolved ~28 hours after onset. Crucially, software orchestration could not reroute around physically dead hardware: stateful volumes on damaged racks required snapshot-restore or launching in unaffected zones, and some resources stayed impaired past resolution.

Blast radius and the single-AZ lesson

Containment held at the AZ boundary by design - impact stayed within one hall of use1-az4 and traffic was shifted away from the zone. Yet load concentration behind a single heavily used zone produced an outsized customer impact, and at least one major customer (Coinbase, ~7h down) failed despite architecture meant to tolerate a single-AZ loss. The lesson is architectural on both sides: operators must not concentrate load behind one hall, and customers must test - not assume - multi-AZ failover.

Sourcing, disclosure gaps, and verification status

AWS published NO formal Post-Event Summary for this event - independently VERIFIED against the AWS PES archive on 2026-08-02 (latest entry: Oct 19 2025 DynamoDB). Consequently there is no disclosed component root cause, no numeric temperature, and no maintenance/inspection history; these must be stated as 'not disclosed', never as 'no lapse occurred'. Equally important: nearly all narrative detail here derives from secondary reporting (SingleStore, Axis Intelligence, AWS Builder Center, Yahoo Finance, Network World) that could NOT be re-fetched during this review (search budget exhausted, no source URLs supplied, guessed URLs 404'd). Those quotations and timestamps are reported-not-confirmed and should be re-verified before publication.

Technical deep-dive

Forensically this incident inverts the classic data-center fire chain and must not be narrated as a fire. No ignition, no aspirating/VESDA smoke detection, no pre-action or clean-agent suppression discharge, no evacuation, and no fire-department arrival are reported anywhere. The detection that mattered was THERMAL/ENVIRONMENTAL: rack- and hall-level temperature sensors registering temperatures crossing safe operating thresholds once cooling capacity collapsed. The reported AWS wording ("temperatures inside a single data center within availability zone use1-az4 exceeded safe operating thresholds") never itemizes the failed mechanical component; the plant is described only generically (chilled-water systems, cooling towers, thermal-management infrastructure).\n\nThe "suppression" analogue was not a fire-suppression system at all; it was the servers' own firmware-level over-temperature protective shutdown. This is the pivotal forensic point: the protection worked as designed and saved silicon, yet its unavoidable side effect was to DE-ENERGISE the affected racks. The reported Health Dashboard line ("EC2 instances and EBS volumes hosted on impacted hardware are affected by the loss of power during the thermal event") makes clear the customer-visible outage was manufactured by a safety mechanism doing its job, not by a mechanism failing.\n\nContainment held at the Availability-Zone boundary by design: impact was confined to a single data hall within use1-az4, and traffic was reportedly shifted away from that zone. The emergency response was an ENGINEERING cooling-restoration effort, not an emergency-services one, and it was deliberately STAGED because a rushed restart could thermally cycle and damage already-stressed equipment - a subtle but correct procedure after a thermal excursion.\n\nThe recovery tail exposes the hard limit of software resilience against physical damage. Because the power loss physically damaged some hardware, software orchestration could not automatically reroute around it - a stateful EBS volume on a dead rack does not migrate itself. Customers were reportedly advised to restore from EBS snapshots or launch in unaffected zones, and some instances/volumes remained impaired even after the incident was marked resolved. The architectural lesson: at least one major customer (Coinbase) went dark despite designs meant to tolerate a single-AZ failure - single-AZ tolerance on paper is not single-AZ tolerance in production. IMPORTANT CAVEAT: every AWS/vendor and press quotation in this deep dive is drawn from secondary reporting that could not be independently re-fetched this review session; treat the specific wording and timestamps as reported-not-confirmed.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home