Azure West US 2 Grid-Disturbance Cooling Loss and Thermal Shutdown (May 2026)
A grid power-quality disturbance reduced cooling capacity across multiple Azure West US 2 availability zones, triggering automated thermal shutdowns of hardware and multi-service impact.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Microsoft Azure
- Data center
- Azure West US 2 (Quincy, Washington)
- Location
- Quincy, Washington, United States
- Date
- 2026-05-01
Impact & scale
- Users affected
- Azure West US 2 customers across multiple AZs
- Financial
- Not published
- Scope
- Grid/cooling thermal event (multi-AZ)
- Azure compute/storage in West US 2 (multi-AZ)
Impact data & metrics
| Total incident duration (impact window) | ~22 hours (04:24 UTC 29 May to 02:30 UTC 30 May 2026) |
| Grid disturbance to customer impact | ~10 minutes (~04:14 UTC storm to 04:24 UTC impact) |
| Cooling loss to automated thermal shutdown | ~10-16 minutes (04:24 impact to ~04:40 shutdown) |
| Cooling restored / temperatures stabilized | ~05:55 UTC 29 May (~1.5 h after impact) |
| Virtual machines recovered | ~50% by 06:15 UTC; ~95% by 12:00 UTC 29 May |
| Storage validation duration | ~14 hours |
| Availability zones affected | 2 physical AZs (AZ-01 and AZ-03) |
| Full mitigation incl. telemetry backlog clearance | 02:30 UTC 30 May 2026 |
Magnitude profile
Grid power-quality disturbance cut cooling across multiple AZs, triggering thermal shutdowns — sub-scores ESTIMATED from public impact reporting, pending deep research.
Sequence of events (SOE)
- TRIGGER Severe thunderstorm with lightning drives 'voltage sag/swell events' across multiple West US 2 (Quincy, WA) datacenter facilities; utility power remains present but unstable. No fire — a grid power-quality disturbance.
- DETECTION Mechanical cooling protection components sense 'abnormal electrical conditions' on the plant. Detection is electrical/thermal, not smoke or flame — no fire-detection system involved.
- IMPACT PIR customer-impact window begins. Mechanical cooling protection components enter a 'lockout' protective state; cooling capacity drops.
- CASCADE Cooling lockout latches off — a 'full lockout rather than staged or degraded operation' — and prevents parts of the cooling system from automatically returning to normal; data-hall temperatures begin rising toward safe-operating thresholds.
- MITIGATION Azure datacenter operations begin manual cooling diagnosis and restoration. No external emergency services — there was no fire; response is Azure's own teams. No evacuation reported (lights-out mechanical/data-hall context).
- CASCADE Protective de-energisation: as temperatures 'rose beyond safe operating thresholds,' cloud infrastructure automatically shuts down 'to prevent damage' and 'preserve data integrity' across two physical availability zones (AZ-01, AZ-03). Equipment self-protection, no agent/water discharge.
- CASCADE Power behavior diverges: some datacenters transfer to on-site generator power while others do not, because utility power 'remained available' despite instability — generator/UPS transfer inconsistent under sag/swell.
- RECOVERY All cooling restored in impacted datacenters and temperatures stabilized (~1.5 h from impact). Physical/thermal hazard contained; the long tail is now data/service recovery, not physical risk.
- RECOVERY Approximately 50% of affected virtual machines recovered as infrastructure is brought back online.
- CASCADE Storage layer requires extended validation before dependent compute can resume; storage validation extended over ~14 hours, constrained by platform limitations and manual, sequential coordination.
- RECOVERY Approximately 95% of VMs recovered after storage largely restored.
- RECOVERY All services recovered except Application Insights and Log Analytics, which are processing telemetry backlogs; Log Analytics could not self-recover from transient initialization failures and over-depended on a single cluster.
- RESTORED Full mitigation declared, including clearing telemetry backlogs (PIR window end). Total incident duration ~22 hours.
- RESTORED Remediation program: cooling recovery processes + Log Analytics startup resiliency (June 2026); networking storage-node speed limits + protective-state tooling improvements (July/August 2026); telemetry multi-cluster architecture (August 2026); holistic cooling fail-state / staged-degradation review (October 2026).
Root cause
Contributing factors
- EXTERNAL TRIGGER: A severe thunderstorm with lightning caused 'utility power disturbances across multiple datacenter facilities' as 'voltage sag/swell events' where 'utility power remained present but was unstable' (GHRP-84G) — a degraded-but-present supply condition that is harder to protect against than a clean outage.
- DESIGN — HARD LOCKOUT, NO STAGED DEGRADATION: Cooling protection 'entered a lockout protective state'; the PIR states the design 'resulted in a full lockout rather than staged or degraded operation' and a holistic cooling fail-state review is committed for October 2026, confirming the all-or-nothing design as a contributing factor.
- DESIGN — NO COOLING AUTO-RECOVERY: The lockout 'prevented parts of the cooling system from automatically returning to normal operation,' so cooling could not self-restore once the transient passed and required manual human diagnosis and restoration.
- COMMISSIONING/TUNING (observed, not confirmed as a lapse): 'Some datacenters transferred to generator power, while others did not, as utility power remained available' — inconsistent generator/UPS transfer under sag/swell implies threshold/transfer tuning sensitivity; the PIR reports the divergence but does not state it as a lapse.
- PROCESS — MANUAL, UNDER-DOCUMENTED COOLING RECOVERY: Recovery depended on manual cooling diagnosis and restoration; Azure committed to improving cooling recovery processes and documentation (target June 2026), indicating the manual process contributed to duration.
- PROCESS/PLATFORM — SLOW STORAGE RETURN-TO-SERVICE: Storage validation extended over ~14 hours; the PIR attributes this to 'platform limitations that constrained the speed at which storage nodes could be returned to service' and to 'dependencies on manual coordination steps' (Networking remediation, July 2026) — a recovery-architecture gap, not a maintenance failure.
- DESIGN — TELEMETRY SINGLE-CLUSTER DEPENDENCY & NO SELF-RECOVERY: Log Analytics could not self-recover from transient initialization failures and over-depended on a single cluster; remediations add startup resiliency (June 2026) and multi-cluster flexibility 'reducing dependence on any single cluster or resource type' (August 2026), making telemetry the longest tail (backlog cleared 02:30 UTC 30 May).
- TOOLING GAP: Weak tooling to 'identify which infrastructure devices have entered protective states and require manual recovery' slowed prioritized recovery across the two affected availability zones (Networking remediation, target August 2026).
- NOTE — NO MAINTENANCE/INSPECTION LAPSE EVIDENCED: The PIR neither evidences nor alleges any maintenance, inspection, or testing failure; root cause is framed as an external grid event plus design/automation gaps. Absence of disclosure is not proof of absence, but none is on the record.
Correction of errors (COE)
- Holistic cooling system fail-state review — evaluate staged/degraded operation vs. full lockout and enable automatic recovery once power stabilizes
- Improve manual cooling recovery processes and documentation to shorten diagnosis/restoration
- Add Log Analytics startup resiliency (self-recover from transient init failures) and multi-cluster flexibility to reduce single-cluster dependence
- Remove platform limitations constraining storage-node return-to-service speed and reduce manual coordination dependencies
- Improve tooling to identify devices in protective states and prioritize foundational infrastructure, with cross-team awareness of thermal events on shared infrastructure
- Review/standardize generator-UPS automatic-transfer thresholds under voltage sag/swell across Quincy facilities
Lessons learnt
- A degraded-but-present grid (sag/swell) is a distinct and harder failure mode than a clean blackout: binary present/absent transfer logic can leave protection in an ambiguous band, producing inconsistent facility responses to one disturbance.
- Protective lockouts should degrade in stages and self-recover: an all-or-nothing latching trip that 'prioritizes equipment protection over operational continuity' converts a transient power-quality event into a multi-hour outage even when no hardware is damaged.
- Equipment self-protection working correctly can still create a large blast radius: automated thermal shutdown across two availability zones prevented hardware damage but caused a broad multi-service outage — protection success and availability are not the same objective.
- Recovery duration was dominated by manual, sequential, single-cluster processes, not physical damage: cooling was restored in ~1.5 h but full service/data recovery took ~22 h, so recovery automation (parallel storage validation, telemetry multi-cluster, protective-state tooling) is where resilience is won.
- Single-source, operator-authored records need explicit confidence framing: the entire account rests on Azure's own PIR with no independent press/regulator/utility corroboration, so figures should be labeled operator-reported.
Improvements & remediation
- SAFETY / COOLING FAIL-STATE (committed, October 2026): Holistic cooling system fail-state review to evaluate staged/degraded operation instead of full lockout and to enable automatic return to normal once power stabilizes — directly targeting the all-or-nothing, non-self-recovering protection that caused the thermal excursion.
- PROCESS — COOLING RECOVERY (committed, June 2026): Improve manual cooling diagnosis/restoration processes and documentation so operators can restore cooling faster after a protective lockout.
- MAINTENANCE / COMMISSIONING (analyst-recommended; not explicitly committed in PIR): Review and standardize generator/UPS automatic-transfer thresholds and protection settings under voltage sag/swell across all Quincy facilities, since 'some datacenters transferred to generator power, while others did not' under an identical disturbance — a commissioning/protection-tuning verification the PIR observes but does not commit to.
- PLATFORM — STORAGE RETURN-TO-SERVICE (committed, July 2026): Remove networking/platform limitations that constrained the speed of returning storage nodes to service and reduce dependence on manual coordination steps.
- RESILIENCE — TELEMETRY (committed, June & August 2026): Add Log Analytics startup resiliency to self-recover from transient initialization failures, and add multi-cluster flexibility to reduce dependence on any single cluster or resource type.
- TOOLING — PROTECTIVE-STATE VISIBILITY (committed, August 2026): Improve tooling/processes to identify which infrastructure devices have entered protective states and require manual recovery, and improve cross-team awareness when thermal events affect shared infrastructure.
Comprehensive analysis
No fire — an electrical-then-thermal chain, not combustion
Despite being catalogued alongside fire/thermal events, GHRP-84G contains no combustion, ignition, smoke, or fire-service element. The verified chain is: external grid sag/swell -> cooling-plant protective lockout -> loss of cooling capacity -> data-hall temperatures beyond safe thresholds -> automated IT shutdown to prevent hardware damage. Every layer is electrical or thermal; the 'fire' framing is a mis-shelving that the record does not support.
The degraded-but-present grid problem
The initiating condition was 'voltage sag/swell events' in which 'utility power remained present but was unstable.' This ambiguous band defeated binary transfer logic: 'some datacenters transferred to generator power, while others did not.' One disturbance, divergent facility responses — a threshold/commissioning sensitivity rather than a single failed component, and the latent weakness that let the disturbance reach cooling-plant protection at all.
Lockout as the pivotal design defect
The PIR is unusually explicit that the cooling design 'resulted in a full lockout rather than staged or degraded operation' and that the lockout 'prevented parts of the cooling system from automatically returning to normal operation.' Two properties — all-or-nothing withdrawal of capacity and latching with no auto-recovery — are what turned a passing power transient into a multi-hour outage. Azure's October-2026 holistic fail-state review targets exactly this.
Where the ~22 hours actually went
Physical containment was fast: cooling restored and temperatures stabilized by ~05:55 UTC, roughly 1.5 hours after impact, with no hardware damage. The remaining ~20 hours were recovery overhead: VMs 50% by 06:15 and 95% by 12:00, storage validation over ~14 hours constrained by platform speed limits and manual coordination, and telemetry (Application Insights / Log Analytics) trailing to 02:30 UTC 30 May due to inability to self-recover and single-cluster dependence. The duration is a recovery-automation story, not a damage story.
Sourcing and confidence
The account rests on one official operator source, Azure PIR GHRP-84G, which was re-fetched live and confirmed on 2026-08-02: tracking ID, 29-30 May 2026 dates, sag/swell and full-lockout wording, AZ-01/AZ-03 naming, 50%/95% VM figures, ~14-hour storage validation, single-cluster telemetry dependency, and June/July/August/October 2026 remediation dates all verify against the live page. The material gap is the absence of independent, non-operator corroboration (press, regulator, Grant County PUD / Bonneville); all figures are operator-reported. The dossier's original 2026-05-01 date is incorrect and is corrected to 29-30 May 2026.
Technical deep-dive
References & provenance
- vendor-status Azure Post Incident Review — Tracking ID GHRP-84G, West US 2 (May 2026)“Some datacenters transferred to generator power, while others did not, as utility power remained available”https://azure.status.microsoft/en-us/status/history/
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).