← All incidents
Incident dossier · Rank #36

Azure West US 2 Grid-Disturbance Cooling Loss and Thermal Shutdown (May 2026)

Microsoft Azure 2026-05-01 8h 0m core impact PowerCooling

A grid power-quality disturbance reduced cooling capacity across multiple Azure West US 2 availability zones, triggering automated thermal shutdowns of hardware and multi-service impact.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Power (2026-05-01)Trigger · Power2026-05-012026-05-01Primary fault at Microsoft Azure — Azure West US 2 (Quincy, Washington)Microsoft AzureAzure West US 2 (Quincy, Washington)Azure West US 2Downstream service degraded by the fault: Azure compute/storage in West US 2 (multi-AZ)Azure compute/storage in West US2…

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Microsoft Azure
Data center
Azure West US 2 (Quincy, Washington)
Location
Quincy, Washington, United States
Date
2026-05-01

Impact & scale

Users affected
Azure West US 2 customers across multiple AZs
Financial
Not published
Scope
Grid/cooling thermal event (multi-AZ)
Services / systems down
  • Azure compute/storage in West US 2 (multi-AZ)

Impact data & metrics

Total incident duration (impact window)~22 hours (04:24 UTC 29 May to 02:30 UTC 30 May 2026)
Grid disturbance to customer impact~10 minutes (~04:14 UTC storm to 04:24 UTC impact)
Cooling loss to automated thermal shutdown~10-16 minutes (04:24 impact to ~04:40 shutdown)
Cooling restored / temperatures stabilized~05:55 UTC 29 May (~1.5 h after impact)
Virtual machines recovered~50% by 06:15 UTC; ~95% by 12:00 UTC 29 May
Storage validation duration~14 hours
Availability zones affected2 physical AZs (AZ-01 and AZ-03)
Full mitigation incl. telemetry backlog clearance02:30 UTC 30 May 2026

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 6Users affected (0–10) — breadth of the user/customer population impacted. — scored 6/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 5Outage duration (0–10) — how long service was degraded/down. — scored 5/10.Blast 6Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 6/10.
Magnitude 5.6 = blast 6×0.35 + users 6×0.25 + financial 5×0.20 + duration 5×0.20 (sub-scores 0–10 · weighted composite)

Grid power-quality disturbance cut cooling across multiple AZs, triggering thermal shutdowns — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events2026-05-29 ~04:14 UTC · TRIGGER — Severe thunderstorm with lightning drives 'voltage sag/swell events' across multiple West US 2 (Quincy, WA) datacenter facilities; utility power remains present but unstable. No fire — a grid power-quality disturbance.TRIGGER2026-05-29 ~04:14 U2026-05-29 ~04:14 UTC · DETECTION — Mechanical cooling protection components sense 'abnormal electrical conditions' on the plant. Detection is electrical/thermal, not smoke or flame — no fire-detection system involved.DETECTION2026-05-29 ~04:14 U2026-05-29 04:24 UTC · IMPACT — PIR customer-impact window begins. Mechanical cooling protection components enter a 'lockout' protective state; cooling capacity drops.IMPACT2026-05-29 04:24 U2026-05-29 ~04:24 UTC · CASCADE — Cooling lockout latches off — a 'full lockout rather than staged or degraded operation' — and prevents parts of the cooling system from automatically returning to normal; data-hall temperatures begin rising toward safe-operating thresholds.CASCADE2026-05-29 ~04:24 U2026-05-29 ~04:31 UTC · MITIGATION — Azure datacenter operations begin manual cooling diagnosis and restoration. No external emergency services — there was no fire; response is Azure's own teams. No evacuation reported (lights-out mechanical/data-hall context).MITIGATION2026-05-29 ~04:31 U2026-05-29 ~04:40 UTC · CASCADE — Protective de-energisation: as temperatures 'rose beyond safe operating thresholds,' cloud infrastructure automatically shuts down 'to prevent damage' and 'preserve data integrity' across two physical availability zones (AZ-01, AZ-03). Equipment self-protection, no agent/water discharge.CASCADE2026-05-29 ~04:40 U2026-05-29 ~04:40 UTC · CASCADE — Power behavior diverges: some datacenters transfer to on-site generator power while others do not, because utility power 'remained available' despite instability — generator/UPS transfer inconsistent under sag/swell.CASCADE2026-05-29 ~04:40 U2026-05-29 ~05:55 UTC · RECOVERY — All cooling restored in impacted datacenters and temperatures stabilized (~1.5 h from impact). Physical/thermal hazard contained; the long tail is now data/service recovery, not physical risk.RECOVERY2026-05-29 ~05:55 U2026-05-29 06:15 UTC · RECOVERY — Approximately 50% of affected virtual machines recovered as infrastructure is brought back online.RECOVERY2026-05-29 06:15 U2026-05-29 ~06:15-12:00 UTC · CASCADE — Storage layer requires extended validation before dependent compute can resume; storage validation extended over ~14 hours, constrained by platform limitations and manual, sequential coordination.CASCADE2026-05-29 ~06:15-12:00 U2026-05-29 12:00 UTC · RECOVERY — Approximately 95% of VMs recovered after storage largely restored.RECOVERY2026-05-29 12:00 U2026-05-29 20:18 UTC · RECOVERY — All services recovered except Application Insights and Log Analytics, which are processing telemetry backlogs; Log Analytics could not self-recover from transient initialization failures and over-depended on a single cluster.RECOVERY2026-05-29 20:18 U2026-05-30 02:30 UTC · RESTORED — Full mitigation declared, including clearing telemetry backlogs (PIR window end). Total incident duration ~22 hours.RESTORED2026-05-30 02:30 U2026-06 to 2026-10 (committed) · RESTORED — Remediation program: cooling recovery processes + Log Analytics startup resiliency (June 2026); networking storage-node speed limits + protective-state tooling improvements (July/August 2026); telemetry multi-cluster architecture (August 2026); holistic cooling fail-state / staged-degradation review (October 2026).RESTORED2026-06 to 2026-

Root cause

There was no fire. This incident is frequently mis-shelved as a thermal/fire event, but the official Post Incident Review (tracking ID GHRP-84G, azure.status.microsoft) describes a strictly electrical-then-thermal chain: an external grid power-quality disturbance drove a protective lockout in the mechanical cooling plant, which in turn forced automated thermal shutdowns of IT infrastructure. No combustion, ignition, smoke, or fire-service involvement is reported anywhere in the record. INITIATING (external) event: A severe thunderstorm producing lightning caused "utility power disturbances across multiple datacenter facilities" serving Azure West US 2 (Quincy, Washington) at approximately 04:14 UTC on 29 May 2026. Crucially these were "voltage sag/swell events" where "utility power remained present but was unstable" — transient under/over-voltage rather than a clean blackout. Because the supply was technically still present, transfer thresholds were not cleanly tripped everywhere: "Some datacenters transferred to generator power, while others did not, as utility power remained available." This inconsistent transfer is the first latent weakness — protection tuning that assumes a binary present/absent supply rather than a degraded/unstable one. PROXIMATE failure mechanism: The unstable voltage was seen by protection components inside the mechanical cooling systems, which "entered a 'lockout' protective state." The PIR states the design "prioritizes equipment protection and data security over operational continuity, which resulted in a full lockout rather than staged or degraded operation," and the lockout "prevented parts of the cooling system from automatically returning to normal operation." This is the core defect: protection latched hard-off with no staged degradation and no auto-recovery, so cooling capacity collapsed and could only be restored by manual human diagnosis. SPECIFIC EQUIPMENT — NOT DISCLOSED (record silent). The PIR does not name the cooling-plant vendor, nor whether the affected components were chillers, CRAH/CRAC units, pumps, or cooling-tower VFD drives, nor the make, model, age, or firmware of the protection relay/controller that latched into lockout. There is no battery or electrolyte-chemistry element because there was no electrical/battery fire. Age, maintenance history, and inspection/testing records are not published. Every equipment-identity claim beyond "a subset of components within the mechanical cooling systems" would be unsupported. CASCADING (thermal) mechanism: With cooling degraded, data-hall temperatures "rose beyond safe operating thresholds," and from ~04:40 UTC cloud infrastructure automatically shut down "to prevent damage" and "preserve data integrity" across two physical availability zones (named AZ-01 and AZ-03 in the PIR). This equipment-protection layer WORKED: no hardware fire or hardware damage is reported. But it converted a cooling fault into a broad multi-service outage. LATENT ROOT (design/automation, not maintenance): Azure's own remediation list frames root cause as an external environmental/grid event compounded by design and automation gaps, not a maintenance, inspection, or testing lapse — none is evidenced or alleged. The latent roots are: (1) cooling protection that chooses full lockout over staged degradation and cannot self-recover; (2) a manual, under-documented cooling recovery process; (3) storage return-to-service bottlenecked on platform limitations and manual, sequential coordination; (4) telemetry (Log Analytics) unable to self-recover from transient init failures and over-dependent on a single cluster; and (5) inconsistent generator/UPS transfer under sag/swell, implying a commissioning/threshold-tuning question that the PIR observes but neither confirms as a lapse nor explains. VERIFICATION STATUS: The account derives from a single official operator source, Azure PIR GHRP-84G. That source was re-fetched live and confirmed on 2026-08-02 — the tracking ID, the 29–30 May 2026 dates, the sag/swell and lockout wording, the AZ-01/AZ-03 naming, the ~22-hour duration, the 50%/95% VM figures, the ~14-hour storage validation, the single-cluster telemetry dependency, and the June/July/August/October 2026 remediation dates all verify verbatim or near-verbatim against the live PIR. The remaining limitation is the absence of independent, non-operator corroboration: no press, regulator, or utility (Grant County PUD / Bonneville) confirmation was obtainable, and figures are operator-reported. Separately, a date-of-record correction applies: this dossier was queued under 2026-05-01, but the official incident is 29–30 May 2026 — the 1 May date is incorrect and is struck.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

No fire — an electrical-then-thermal chain, not combustion

Despite being catalogued alongside fire/thermal events, GHRP-84G contains no combustion, ignition, smoke, or fire-service element. The verified chain is: external grid sag/swell -> cooling-plant protective lockout -> loss of cooling capacity -> data-hall temperatures beyond safe thresholds -> automated IT shutdown to prevent hardware damage. Every layer is electrical or thermal; the 'fire' framing is a mis-shelving that the record does not support.

The degraded-but-present grid problem

The initiating condition was 'voltage sag/swell events' in which 'utility power remained present but was unstable.' This ambiguous band defeated binary transfer logic: 'some datacenters transferred to generator power, while others did not.' One disturbance, divergent facility responses — a threshold/commissioning sensitivity rather than a single failed component, and the latent weakness that let the disturbance reach cooling-plant protection at all.

Lockout as the pivotal design defect

The PIR is unusually explicit that the cooling design 'resulted in a full lockout rather than staged or degraded operation' and that the lockout 'prevented parts of the cooling system from automatically returning to normal operation.' Two properties — all-or-nothing withdrawal of capacity and latching with no auto-recovery — are what turned a passing power transient into a multi-hour outage. Azure's October-2026 holistic fail-state review targets exactly this.

Where the ~22 hours actually went

Physical containment was fast: cooling restored and temperatures stabilized by ~05:55 UTC, roughly 1.5 hours after impact, with no hardware damage. The remaining ~20 hours were recovery overhead: VMs 50% by 06:15 and 95% by 12:00, storage validation over ~14 hours constrained by platform speed limits and manual coordination, and telemetry (Application Insights / Log Analytics) trailing to 02:30 UTC 30 May due to inability to self-recover and single-cluster dependence. The duration is a recovery-automation story, not a damage story.

Sourcing and confidence

The account rests on one official operator source, Azure PIR GHRP-84G, which was re-fetched live and confirmed on 2026-08-02: tracking ID, 29-30 May 2026 dates, sag/swell and full-lockout wording, AZ-01/AZ-03 naming, 50%/95% VM figures, ~14-hour storage validation, single-cluster telemetry dependency, and June/July/August/October 2026 remediation dates all verify against the live page. The material gap is the absence of independent, non-operator corroboration (press, regulator, Grant County PUD / Bonneville); all figures are operator-reported. The dossier's original 2026-05-01 date is incorrect and is corrected to 29-30 May 2026.

Technical deep-dive

Fault propagation, layer by layer, as disclosed in GHRP-84G (re-verified against the live PIR on 2026-08-02). Layer 0 — the grid. A lightning-producing thunderstorm injected "voltage sag/swell events" into the utility feeds of multiple Quincy datacenter facilities at ~04:14 UTC. The defining characteristic is that "utility power remained present but was unstable." Datacenter power architecture is generally engineered around a binary decision — utility good vs. utility lost — with automatic transfer switches and UPS/generator start logic keyed to voltage/frequency windows. A sag/swell that dips and recovers repeatedly without fully dropping out can sit inside the ambiguous band: enough to disturb sensitive downstream protection, not enough to cleanly command a transfer. The observed symptom confirms this: "Some datacenters transferred to generator power, while others did not." Identical disturbance, divergent responses — a signature of threshold/commissioning sensitivity rather than a single failed component. Layer 1 — cooling-plant protection. The mechanical cooling systems contain their own electrical protection (motor protection relays, VFD DC-bus/under-voltage trips, or controller interlocks — the specific device is NOT disclosed). Per the PIR, a subset of components "entered a 'lockout' protective state." Two design choices compound. First, lockout is all-or-nothing: capacity was not throttled or staged; the PIR is explicit that the design "resulted in a full lockout rather than staged or degraded operation." Second, lockout is latching: it "prevented parts of the cooling system from automatically returning to normal operation." Once the transient passed and voltage was nominal again, the plant did not restart itself. This is precisely the behavior Azure's remediation targets — a "holistic cooling system fail-state review" is committed for October 2026. Layer 2 — thermal. Cooling capacity dropped while IT load — and therefore heat output — continued. Data-hall temperatures climbed until they "rose beyond safe operating thresholds." The PIR gives no numeric temperature, threshold, or rate-of-rise; those figures are not disclosed. Time-to-threshold is inferable only indirectly: impact began 04:24, and automated infrastructure shutdown engaged from ~04:40 — on the order of 10–16 minutes from cooling loss to thermal-shutdown threshold, consistent with high-density halls having little thermal ride-through once cooling latches off. Layer 3 — IT self-protection. To avoid hardware damage, cloud infrastructure shut down "to prevent damage" and "preserve data integrity" across two physical availability zones (AZ-01, AZ-03). This is the point at which a mechanical fault became a customer-facing multi-service outage. The protection succeeded on its own terms — no hardware damage is reported — but a two-AZ automated shutdown is a large blast radius. Layer 4 — the recovery tail, which is where most of the ~22-hour duration lives. Cooling itself was restored and temperatures stabilized by ~05:55 UTC (~1.5 h from impact) — the physical hazard was contained quickly. But bringing infrastructure back was sequential and manual: ~50% of VMs recovered by 06:15, ~95% only by 12:00, with storage validation extending over ~14 hours. Storage must be integrity-checked before dependent compute can safely resume; the PIR attributes the slow storage-node return to "platform limitations that constrained the speed" and to "dependencies on manual coordination steps" (Networking-team remediation, July 2026). The longest tail was telemetry: by 20:18 UTC all services had recovered except Application Insights and Log Analytics, which were processing backlogs and could not self-recover from transient initialization failures; full mitigation including backlog clearance was declared 02:30 UTC on 30 May. Log Analytics' single-cluster dependency (remediation "reducing dependence on any single cluster or resource type's availability," multi-cluster flexibility, August 2026) made it the slowest component. The forensic shape is an inverted pyramid: a sub-second-to-minutes grid transient at the top, a ~10–16-minute thermal excursion, a ~1.5-hour cooling restoration, and a ~22-hour service/data recovery driven almost entirely by manual, sequential, single-cluster recovery processes rather than by any physical damage.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home