AWS US-EAST-1 Cooling/Thermal Event and use1-az4 Power Loss
On 7 May 2026 cooling systems failed in AWS's use1-az4 availability zone in US-EAST-1 (Northern Virginia). As data-hall temperatures crossed operating thresholds, servers automatically shut down to protect hardware, cutting power to the EC2 instances and EBS volumes hosted on that gear. The compute and storage loss cascaded to dependent services — IoT Core, ELB, NAT Gateway and Redshift — with elevated error rates and slower provisioning. AWS shifted traffic off the impaired zone and advised customers to relaunch or redistribute workloads in other US-EAST-1 availability zones. AWS first flagged the problem at 5:25 PM PDT on 7 May and restored cooling to pre-incident capacity at 1:50 PM PDT on 8 May, roughly 20 hours 25 minutes later. Downstream, Coinbase reported a multi-hour outage in which users could not trade or move funds. No formal AWS post-incident report was published in the available sources.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Amazon Web Services
- Data center
- US-EAST-1 (use1-az4), Northern Virginia
- Location
- Ashburn, United States, use1-az4
- Date
- 2026-05-07
Impact & scale
- Users affected
- Not publicly quantified. Directly impaired EC2/EBS customers in use1-az4 plus downstream users of dependent services; named downstream victim Coinbase suffered a multi-hour trading and funds-movement outage, but no affected-user count was disclosed by any source.
- Financial
- Not disclosed. No source states a dollar loss for AWS or for downstream customers; Coinbase impact was reported only qualitatively.
- Scope
- Major single-AZ regional impairment with multi-service cascade and named downstream customer outage
- EC2 (direct)
- EBS (direct)
- IoT Core (cascaded)
- ELB (cascaded)
- NAT Gateway (cascaded)
- Redshift (cascaded)
Impact data & metrics
| Total incident duration (onset to resolved) | ~28 hours (reported) |
| Cooling-restoration duration (onset to pre-incident capacity) | ~21.5 hours (4:20 PM PDT May 7 -> 1:50 PM PDT May 8) |
| Threshold-exceedance / trigger time | 4:20 PM PDT May 7 (23:20 UTC; Axis separately cites 23:50 UTC - internal inconsistency) |
| Coinbase customer downtime | ~7 hours (reported) |
| Blast radius | 1 data hall / 1 Availability Zone (use1-az4), described as a very heavily used AWS zone |
| Cooling units failed | 'Multiple' / 'multiple chiller units' near-simultaneously (exact count not disclosed) |
| Official AWS post-event summary published | 0 (none for May 2026; latest PES entry is the Oct 19, 2025 DynamoDB disruption) |
| Temperature detail disclosed (deg / setpoint / threshold value) | Not disclosed - only 'exceeded safe operating thresholds', no numeric value |
| Chiller/CRAH make, model, age, refrigerant | Not disclosed in any source |
| Recency ranking of the outage | Reported as 4th significant us-east-1 outage since October 2025 (single-source, unverified) |
Magnitude profile
Scores reflect a long (~20h) single-AZ event whose direct footprint (EC2/EBS in one zone) fanned out to four cascaded AWS services and at least one high-profile downstream customer (Coinbase). User and financial scores are estimates constrained by disclosure: no affected-user count or dollar loss was published, so financial (6) and users (7) are inferred from the scale of the region and named victims rather than measured. Duration (7) is anchored to the two firmly stamped endpoints. Blast radius (6) is moderated by the impairment being confined to one availability zone, not the whole region.
Sequence of events (SOE)
- TRIGGER Multiple cooling/chiller units in a single data hall of Availability Zone use1-az4 (Ashburn) reportedly fail near-simultaneously; hall cooling capacity drops sharply.
- DETECTION Rack- and hall-level temperature sensors register temperatures exceeding safe operating thresholds; detection was thermal/environmental, NOT smoke/fire (no VESDA, no fire alarm reported).
- MITIGATION Servers execute their built-in automatic over-temperature protective shutdown to protect silicon; no fire-suppression system involved.
- IMPACT The protective shutdown de-energises the affected racks; EC2 instances and EBS volumes lose power, some physically damaged.
- MITIGATION AWS reportedly shifts customer traffic away from use1-az4 to contain the blast radius to the single affected zone.
- DETECTION AWS Health Dashboard reportedly posts first status message: 'EC2 instances and EBS volumes hosted on impacted hardware are affected by the loss of power during the thermal event.'
- IMPACT Downstream customers begin outages; Coinbase and other services on affected racks reportedly go dark.
- CASCADE AWS Health Dashboard reportedly warns dependent services will be hit: other AWS services depending on affected EC2/EBS in the AZ 'may also experience impairments.'
- CASCADE Software-orchestration recovery proves impossible for physically damaged hardware: 'Software orchestration cannot automatically reroute around physical hardware damage.'
- RECOVERY AWS reports staged cooling-restoration underway ('incremental progress to restore cooling systems'); recovery deliberately paced to avoid damaging thermally-stressed equipment.
- DETECTION AWS reportedly posts root-cause status message confirming the thermal->shutdown->power-loss chain ('servers automatically shut down ... to protect the hardware').
- RECOVERY AWS advises customers to restore from EBS snapshots or launch resources in unaffected zones, acknowledging some hardware will not come back in place.
- RESTORED Cooling reportedly restored to pre-incident capacity in the affected hall.
- RESTORED Coinbase service reportedly restored after roughly 7 hours of downtime.
- RESTORED Incident marked resolved (~28 hours after onset); however some EC2 instances/EBS volumes reportedly remained impaired due to physical damage.
- RESTORED VERIFIED: AWS Post-Event Summaries archive checked; most recent entry is the Oct 19 2025 DynamoDB disruption. No formal post-event summary for this May 2026 thermal event was ever published.
Root cause
Contributing factors
- Insufficient cooling redundancy within the affected hall: multiple cooling/chiller units reportedly failed near-simultaneously and remaining capacity could not hold the thermal load before thresholds were breached (SingleStore; Axis Intelligence) - secondary sourcing, unverified this session.
- Probable common-cause failure implied but unconfirmed: 'simultaneously' points to a shared dependency (chilled-water loop, shared power feed to the cooling plant, or a BMS/controls fault), yet no source identifies the shared element - a single-point-of-failure contributor AWS never disclosed.
- Extreme load concentration in a single zone: use1-az4 is described as one of the most heavily used AWS zones, so one physical hall's cooling loss translated into an outsized customer blast radius (Axis - unverified).
- Maintenance/inspection/testing status NOT DISCLOSED: AWS published no post-event summary (VERIFIED via the PES archive, latest entry Oct 19 2025 DynamoDB), so whether deferred maintenance, a failed PM action, or a controls misconfiguration contributed is genuinely unknown - not a finding of 'no lapse'.
- Protective-shutdown design coupling: the servers' over-temperature protective shutdown, while correct, de-energised racks and physically damaged some hardware, converting a thermal event into a power-loss-plus-hardware-damage event (reported AWS Builder Center wording).
- Analyst-hypothesized load growth outpacing legacy cooling: trade commentary raised concern that dense AI/HPC racks strain legacy cooling infrastructure (Axis) - explicitly analyst speculation, not an AWS finding.
- Customer-side single-AZ dependency: despite designs meant to tolerate one-AZ failure, at least one major customer (Coinbase) did not survive the loss of use1-az4 (SingleStore - unverified).
Correction of errors (COE)
- Publish a formal Post-Event Summary identifying the failed cooling component, the common-cause mechanism, temperatures reached, and corrective actions
- Increase in-hall cooling redundancy to survive simultaneous loss of multiple units
- Eliminate shared single points of failure (chilled-water loop, cooling-plant power feed, BMS/controls) enabling common-cause failure
- Disclose and strengthen preventive-maintenance / inspection regime for cooling units with concurrent maintainability
- Formalize staged thermal-recovery runbook and pre-stage spare hardware to shorten physical-damage recovery
- Validate genuine multi-AZ/region failover via regular game-days
Lessons learnt
- A cooling/thermal event can be as destructive as a power event: loss of cooling triggered protective server shutdowns that de-energised racks AND physically damaged hardware - thermal risk must be modelled as a power-and-hardware-loss risk, not just an availability blip.
- A correctly functioning safety mechanism can manufacture the outage: the over-temperature protective shutdown did its job (saved silicon) but its side effect (removing rack power) was the customer-visible failure - safety-shutdown blast radius must be designed for.
- Software resilience cannot reroute around physically dead hardware: stateful EBS volumes on damaged racks did not migrate; snapshot/restore and cross-zone launch were the only paths - stateful recovery design is essential.
- Paper single-AZ tolerance is not production single-AZ tolerance: a major customer went dark despite designs meant to survive one-AZ loss - failover must be tested, not assumed.
- Load concentration amplifies single-hall failures: concentrating enormous customer load behind one heavily used zone/hall turns a localized mechanical fault into a broad outage.
- Absence of an official post-event summary limits forensic certainty: with no AWS PES, component-level root cause, temperatures, and maintenance history remain unknown - the industry cannot fully learn from the event.
Improvements & remediation
- DESIGN / REDUNDANCY: Increase cooling redundancy within each data hall (true N+ concurrent-maintainable capacity) so the loss of multiple units cannot breach thermal thresholds before make-up cooling engages.
- DESIGN / COMMON-CAUSE ELIMINATION: Audit for and remove shared single points of failure across the cooling plant (chilled-water loop segmentation, independent per-unit power feeds, diverse BMS/controls) so multiple units cannot fail together from one root.
- MaintenanceEstablish and disclose a preventive-maintenance and inspection regime for chillers/CRAH/CDU units (vibration, refrigerant, controls-firmware, coolant chemistry) with concurrent-maintainability so PM work never reduces redundancy below survivable levels.
- SAFETY / THERMAL RUNBOOK: Formalize a staged thermal-recovery runbook (controlled re-cooling and phased re-energisation) to prevent thermal-cycling damage to stressed hardware, and pre-stage spare hardware to shorten physical-damage recovery.
- TRANSPARENCY: Publish a formal Post-Event Summary (root cause, failed component, timeline, corrective actions) - none was issued for this event, limiting industry learning and customer trust.
- CUSTOMER ARCHITECTURE: Guide and validate genuine multi-AZ / multi-region failover with regular game-day testing, since paper single-AZ tolerance (e.g. Coinbase) failed in production.
Comprehensive analysis
What actually failed - and what did not
This was a cooling/thermal-management failure in a single data hall of use1-az4, NOT a fire. No ignition, smoke detection, suppression discharge, evacuation, or fire-department response is reported anywhere. Multiple cooling/chiller units reportedly failed near-simultaneously; the mechanical component, make, model, age, refrigerant, and specific fault were never disclosed. 'Simultaneously' implies a common-cause failure (shared coolant loop, shared power feed, or a controls fault) but none is confirmed. All component-level detail remains a disclosed gap.
Why a safety mechanism manufactured the outage
When cooling collapsed, rack-inlet temperatures crossed safe operating thresholds and servers executed their built-in over-temperature protective shutdown to protect silicon. That shutdown, working exactly as designed, de-energised the affected racks - and the power loss physically damaged some EC2 instances and EBS volumes. The customer-visible outage was therefore produced by protection succeeding, not failing. This is the pivotal, counter-intuitive forensic point and is the correct frame for the whole incident.
The physical-damage recovery tail
Recovery was an engineering cooling-restoration effort, deliberately staged because a rushed re-energisation could thermally cycle and damage already-stressed hardware. Cooling was reportedly restored to pre-incident capacity ~21.5 hours after onset, with the incident resolved ~28 hours after onset. Crucially, software orchestration could not reroute around physically dead hardware: stateful volumes on damaged racks required snapshot-restore or launching in unaffected zones, and some resources stayed impaired past resolution.
Blast radius and the single-AZ lesson
Containment held at the AZ boundary by design - impact stayed within one hall of use1-az4 and traffic was shifted away from the zone. Yet load concentration behind a single heavily used zone produced an outsized customer impact, and at least one major customer (Coinbase, ~7h down) failed despite architecture meant to tolerate a single-AZ loss. The lesson is architectural on both sides: operators must not concentrate load behind one hall, and customers must test - not assume - multi-AZ failover.
Sourcing, disclosure gaps, and verification status
AWS published NO formal Post-Event Summary for this event - independently VERIFIED against the AWS PES archive on 2026-08-02 (latest entry: Oct 19 2025 DynamoDB). Consequently there is no disclosed component root cause, no numeric temperature, and no maintenance/inspection history; these must be stated as 'not disclosed', never as 'no lapse occurred'. Equally important: nearly all narrative detail here derives from secondary reporting (SingleStore, Axis Intelligence, AWS Builder Center, Yahoo Finance, Network World) that could NOT be re-fetched during this review (search budget exhausted, no source URLs supplied, guessed URLs 404'd). Those quotations and timestamps are reported-not-confirmed and should be re-verified before publication.
Technical deep-dive
References & provenance
- official-postmortem AWS Post-Event Summaries (official archive)“Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region - October 19, 2025 (most recent entry; no May 2026 thermal-event summary listed).”https://aws.amazon.com/premiumsupport/technology/pes/
- press The 10 Biggest Cloud Outages Of 2026 (So Far) (CRN)“AWS confirmed the outage was due to overheating caused by a cooling failure at its North Virginia data center, US-East-1, which is a heavily used region.”https://www.crn.com/news/cloud/2026/the-10-biggest-cloud-outages-of-2026-so-far
- news AWS hit by US-East-1 outage after data center thermal event“After cooling systems failed in the affected availability zone, servers automatically shut down when the temperatures exceeded the operating thresholds in order to protect the hardware.”https://www.networkworld.com/article/4168878/aws-hit-by-us-east-1-outage-after-data-center-thermal-event.html
- news AWS warns of EC2 impairment as power loss hits notorious US-East-1 region“EC2 instances and EBS volumes hosted on impacted hardware are affected by the loss of power during the thermal event.”https://www.theregister.com/off-prem/2026/05/08/aws-warns-of-ec2-impairment-as-power-loss-hits-notorious-us-east-1-region/5235509
- postmortem-summary Coinbase publishes postmortem on AWS-triggered outage“Internal Coinbase systems affected: Raft-based matching engine; Kafka event-streaming infrastructure; Order routing services.”https://www.infoq.com/news/2026/06/coinbase-aws-failure-postmortem/
- news AWS outage traced to cooling failure at key Virginia data center“AWS outage traced to cooling failure at key Virginia data center.”https://www.msn.com/en-us/news/insight/aws-outage-traced-to-cooling-failure-at-key-virginia-data-center/gm-GM72630868
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.