← All incidents
Incident dossier · Rank #26

Azure South Central US: Lightning-Induced Cooling Loss and Automated Integrity Shutdown (September 2018)

Microsoft Azure 2018-09-04 38h 20m core impact CoolingPower

On 4 September 2018, a high-energy thunderstorm with lightning struck southern Texas near Microsoft's South Central US datacenters. Per The Register's account of Azure's RCA, a lightning strike at 08:42 UTC switched one datacenter to generator power and overloaded surge suppressors on the mechanical cooling plant, shutting it down; the official VSTS postmortem describes voltage sags and swells on the utility feed that impacted cooling. As temperatures climbed, automated datacenter integrity procedures forced a structured power-down of servers and storage to protect hardware. Utility power was restored in roughly nine hours (per Data Center Dynamics), but recovery of services took far longer: VSTS primary recovery ran about twenty-one hours, the official VSTS incident spanned 09:45 UTC on 4 September to 00:05 UTC on 6 September (the longest outage in that service's history), and broader Azure — on the order of forty services per contemporaneous reporting — was not declared fully mitigated until 08:40 UTC on 7 September. A separate ~2-hour Release Management incident (a database going offline) occurred during recovery.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Cooling (2018-09-04)Trigger · Cooling2018-09-042018-09-04Primary fault at Microsoft Azure — South Central US region (San Antonio-area datacenters)Microsoft AzureSouth Central US region (San Antonio-area datacenters)South Central US regionDownstream service degraded by the fault: All VSTS Scale Units in South Central USAll VSTS Scale Units in SouthDownstream service degraded by the fault: Azure DevOps Marketplace (global)Azure DevOps MarketplaceDownstream service degraded by the fault: Release Management (US users)Release ManagementDownstream service degraded by the fault: Git (intermittent failures)Git+6 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Microsoft Azure
Data center
South Central US region (San Antonio-area datacenters)
Location
San Antonio, United States, South Central US
Date
2018-09-04

Impact & scale

Users affected
Undisclosed by Microsoft; region-scoped impact to South Central US customers plus global Azure DevOps (VSTS) Marketplace users. No user/customer count was published in the official postmortem.
Financial
Not publicly disclosed; no financial figure appears in Microsoft's public write-up.
Scope
Sev 1 / major regional outage (characterized by Microsoft as the longest VSTS outage in the service's history)
Services / systems down
  • All VSTS Scale Units in South Central US
  • Azure DevOps Marketplace (global)
  • Release Management (US users)
  • Git (intermittent failures)
  • Package Management (intermittent failures)
  • VS / VS Code extension acquisition
  • Dashboards
  • User profiles hosted in South Central US
  • Hosted macOS build/release queue
  • ~40 broader Azure services (per contemporaneous public reporting, not enumerated in the official postmortem)

Impact data & metrics

Lightning strike (per Azure RCA via The Register)2018-09-04 08:42 UTC
Incident start (VSTS window)2018-09-04 09:45 UTC
Incident end (VSTS window)2018-09-06 00:05 UTC
VSTS incident duration (official window)~38h20m (2300 min)
VSTS primary recovery~21 hours
Utility power restoration~9 hours after onset (single-source; Data Center Dynamics, not independently re-verified)
Broader Azure recoveryMajority restored 11:00 UTC 5 Sep; full mitigation 08:40 UTC 7 Sep
Azure services affected~40 (contemporaneous public reporting; not enumerated in official postmortem)
Committed remediation (COE) items7
Severity framingLongest VSTS outage in the service's history
Secondary Release Management incident~2 hours (database offline, during recovery)
Published user-impact / damage / temperature / cost figuresNone (confirmed absent from official disclosure)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 7Users affected (0–10) — breadth of the user/customer population impacted. — scored 7/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 8Outage duration (0–10) — how long service was degraded/down. — scored 8/10.Blast 7Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 7/10.
Magnitude 7.0 = blast 7×0.35 + users 7×0.25 + financial 6×0.20 + duration 8×0.20 (sub-scores 0–10 · weighted composite)

Duration and blast-radius scores dominate: a single region's cooling loss cascaded into a multi-service, multi-hour outage with global reach (Marketplace) and the longest VSTS outage on record. User and financial scores are estimated as moderate-high because Microsoft never published customer counts, damaged-hardware counts, or dollar impact; those figures are confirmed absent from the public disclosure and are not inferred here.

Sequence of events (SOE)

Phased sequence of events2018-09-04 early morning UTC · TRIGGER — High-energy storms hit southern Texas near Azure's South Central US region; multiple regional datacenters see voltage sags and swells across the utility feeds.TRIGGER2018-09-04 early morning U2018-09-04 08:42 UTC · TRIGGER — Lightning causes electrical activity on the utility supply, producing significant voltage swells on the feeds serving the region.TRIGGER2018-09-04 08:42 U2018-09-04 08:42 UTC · MITIGATION — Protective transfer: the voltage swells trigger a portion of one Azure datacenter to transfer from utility power to generator power (intended response to a disturbed feed).MITIGATION2018-09-04 08:42 U2018-09-04 08:42 UTC · CASCADE — The same power swells shut down the datacenter's mechanical cooling systems despite surge suppressors being in place — cooling is lost (The Register: suppressors 'overloaded').CASCADE2018-09-04 08:42 U2018-09-04 08:42+ UTC · DETECTION — A load-dependent thermal buffer engineered into the cooling system absorbs heat, temporarily maintaining operational temperatures (finite ride-through, no active cooling).DETECTION2018-09-04 08:42+ U2018-09-04 (time not disclosed) UTC · DETECTION — Thermal buffer depletes; environmental/thermal monitoring registers datacenter temperature exceeding safe operational thresholds (actual temperatures/thresholds not disclosed in RCA).DETECTION2018-09-04 (time not disclosed) U2018-09-04 (time not disclosed) UTC · MITIGATION — Automated integrity shutdown of devices is initiated to preserve infrastructure and data integrity (the functional protective action — not fire suppression; no suppression system was involved).MITIGATION2018-09-04 (time not disclosed) U2018-09-04 ~09:29 UTC · IMPACT — Temperatures rise so fast in parts of the hall that some hardware is damaged before it can shut down: a significant number of storage servers, plus a small number of network devices and power units, are damaged.IMPACT2018-09-04 ~09:29 U2018-09-04 09:29 UTC · IMPACT — Customer impact begins as storage servers shut down on unsafe temperatures, affecting multiple Azure services dependent on those storage servers.IMPACT2018-09-04 09:29 U2018-09-04 (during active storms) UTC · MITIGATION — Onsite engineering teams — working while storms remain active, with no evacuation and no emergency-services dispatch reported — transfer the rest of the datacenter to generators, stabilizing the power supply.MITIGATION2018-09-04 (during active storms) U2018-09-04 (post-shutdown) UTC · CASCADE — Single-region dependency propagates the outage: Azure Service Manager (primary site South Central US, no automatic failover) produces the widest impact for customers outside the region.CASCADE2018-09-04 (post-shutdown) U2018-09-04 (post-shutdown) UTC · CASCADE — A North America Azure AD site in the affected datacenter goes down; authentication traffic begins automatically routing to other sites (self-healing failover).CASCADE2018-09-04 (post-shutdown) U2018-09-04 ~12:30 UTC · CASCADE — Azure status page fails to scale under load due to incorrect auto-scale configuration, returning intermittent 500 errors to a subset of visitors.CASCADE2018-09-04 ~12:30 U2018-09-04 (decision point) UTC · RECOVERY — Recover-not-failover decision: failover is rejected because asynchronous geo-replication would cause limited data loss; teams commit to recovering data in place.RECOVERY2018-09-04 (decision point) U2018-09-04 ~14:40 UTC · RESTORED — Azure AD customer impact mitigated as authentication routing to healthy sites stabilizes.RESTORED2018-09-04 ~14:40 U2018-09-04 (recovery phase 1) UTC · RECOVERY — Recovery step one: restore the Azure Software Load Balancers (SLBs) for the storage scale units.RECOVERY2018-09-04 (recovery phase 1) U2018-09-04 (recovery phase 2) UTC · RECOVERY — Recovery step two: recover storage servers and data — replace failed components, migrate customer data from damaged to healthy servers, and validate that no recovered data is corrupted.RECOVERY2018-09-04 (recovery phase 2) U2018-09-04 ~23:05 UTC · RESTORED — Azure status page impact mitigated after the auto-scale/load issue is resolved.RESTORED2018-09-04 ~23:05 U2018-09-05 01:10 UTC · RESTORED — ASM impact mitigated once the associated South Central US storage servers are brought back online — the out-of-region blast radius closes.RESTORED2018-09-05 01:10 U

Root cause

On September 4, 2018, high-energy thunderstorms struck southern Texas near Microsoft Azure's South Central US region, and multiple datacenters in the region experienced voltage sags and swells on their utility feeds. At approximately 08:42 UTC a lightning-associated electrical transient (a large voltage swell immediately followed by a large sag) hit one datacenter: it switched to generator power and, because the sag fell below the required voltage specification, the chiller plant's protective logic powered the chillers down and locked them out to protect the mechanical equipment. This chiller lock-out was a by-design protective response, not a fault in itself. (The Register, quoting Microsoft's RCA, describes the same event as switching the datacenter to generator power and overloading suppressors on the mechanical cooling system, shutting it down.)\n\nThe proximate root cause of customer impact was not the lightning strike but the failure of cooling redundancy to take over. Under the standard design, the Mechanical/Electrical/Plumbing (MEP) management control system automatically invokes a redundant cooling system until the chiller plant recovers — automation that had withstood similar events in many Microsoft datacenters worldwide. In this specific datacenter, per Microsoft's RCA, that automatic failover had been switched to MANUAL mode following installation of new equipment for which all testing had not yet been completed. With the automatic path disabled, the only remaining safeguard was a documented manual failover that required an engineer to act on cooling alerts. Onsite engineers were simultaneously investigating multiple concurrent alarms, and the cooling alerts were not manually actioned in time. With neither automatic nor timely manual failover, datacenter temperatures rose unchecked once thermal buffers were depleted.\n\nAs temperatures climbed, infrastructure devices and servers hit thermal-integrity shutdown thresholds; but in parts of the datacenter temperatures rose so quickly that a number of storage servers and a small number of network devices were physically damaged rather than shut down cleanly (corroborated by The Register: 'storage units and network devices, were damaged'). Recovery was deliberately prolonged by a data-integrity decision: per Microsoft's RCA, Microsoft chose to recover data in place rather than fail storage over to another region, because the asynchronous nature of geo-replication would have caused limited data loss. Recovery required replacing failed hardware, migrating customer data from damaged to healthy servers, and validating integrity across many servers with substantial manual intervention, extending South Central US impact over multiple days.\n\nThe blast radius extended well beyond South Central US because of resiliency gaps in global control-plane services whose primary metadata site is South Central US. Azure Service Manager (ASM), which handles management operations for 'classic' resource types, is a global service but keeps its primary metadata site in South Central US and does not support automatic failover, so classic management operations were impaired far beyond the region (corroborated by The Register). Azure Active Directory then degraded: per Microsoft's RCA an AAD scale-ahead autoscale operation depended on the ASM API and could not complete, QoS throttling engaged as sites neared safe-utilization limits (producing authentication failures/timeouts — the throttling and timeouts are corroborated by The Register), and an Office-client aggressive-retry bug amplified load. Microsoft's RCA also states the Azure status page itself, with non-optimized auto-scale, returned intermittent HTTP 500 errors during the traffic surge.\n\nSUMMARY: the lightning-induced voltage sag was the trigger; the chiller lock-out was a by-design protective response; the true root cause was that redundant cooling failover had been left in manual mode (pending completion of testing after new-equipment installation) and the corresponding cooling alerts were not manually actioned amid competing alarms, letting temperatures rise fast enough to physically damage storage and network hardware. Insufficient resiliency and lack of automatic failover in dependent global services (ASM, AAD) and the status page then turned a single-datacenter cooling failure into a broad, multi-service incident.\n\nSOURCING NOTE: Microsoft published an official RCA for this incident on its Azure status-history page; in this verification pass that original page could not be re-fetched (archive.org blocked to the fetcher, live status page rotated to current incidents), so the account rests on Microsoft's RCA as relayed and directly corroborated on its central mechanism by The Register. Precise items still resting solely on Microsoft's RCA (manual-mode cause, the 30-minutes-earlier sibling-datacenter manual failover, the Sept-7 08:40 UTC full-mitigation endpoint, the AAD/ASM-API autoscale chain, the Office-client retry bug, and the status-page 500s from ~12:30 UTC) should be treated as officially-stated-but-not-independently-re-verified here.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

This was a thermal-electrical incident, not a fire

The schema's fire-oriented fields (ignition equipment make/model/age/chemistry, fire suppression, evacuation, emergency-services de-energisation) have no basis in the sourced record. There was no combustion and no ignition source. The physical trigger was lightning-induced voltage swells; the damage mechanism was heat after loss of mechanical cooling. Where the schema asks about fire suppression, the functional analogue is the automated integrity shutdown of devices; where it asks about evacuation/emergency response, the record shows none — onsite teams instead stabilized power while storms were still active. These gaps are disclosed rather than filled with fabricated detail.

The electrical-to-thermal cascade

A lightning strike at 08:42 UTC produced voltage swells that both transferred part of one datacenter to generator power (working as intended) and shut down the mechanical cooling systems despite surge suppressors being in place. A load-dependent thermal buffer bridged a finite window, then temperatures exceeded safe thresholds. An automated device shutdown was initiated to protect infrastructure and data, but temperatures rose so fast in some zones that a significant number of storage servers plus a small number of network devices and power units were damaged before they could shut down.

Blast-radius amplification via single-region dependencies

Physical damage was confined to one datacenter, yet the outage spread far via single-region service dependencies. Azure Service Manager's primary metadata site is South Central US and it does not support automatic failover, producing the widest out-of-region impact until associated storage came back at 01:10 UTC on 5 September. A North America Azure AD site and VSTS organizations co-located there added to the impact — with AAD self-healing by routing authentication to other sites (~14:40 UTC).

Recovery strategy and the data-integrity tradeoff

Microsoft deliberately chose recovery over failover: asynchronous geo-replication meant a failover would have caused limited data loss. Recovery was staged — restore the storage Software Load Balancers, then recover storage servers and data by replacing failed components, migrating customer data to healthy servers, and validating no corruption. This prioritized durability over speed and lengthened the outage, a defensible but costly tradeoff rooted in the replication design.

The communications failure compounded the event

Independently of the datacenter, the Azure status page failed to scale under load because of an incorrect auto-scale configuration, returning intermittent 500 errors from roughly 12:30 to 23:05 UTC — precisely when customers needed status information. This is a communications-resiliency defect: the status channel was neither correctly scaled nor sufficiently decoupled from the platform whose health it reports.

Technical deep-dive

Physical layer — the electrical transient. The event began with a lightning strike at 08:42 UTC on 4 September 2018 during high-energy storms over southern Texas. Multiple Azure datacenters in the South Central US region saw voltage sags and swells across the utility feeds. Two consequences followed at the affected datacenter: (a) a portion of it automatically transferred from utility to generator power — the intended protective response to a disturbed feed — and (b) crucially, the same swells shut down the datacenter's mechanical cooling systems despite surge suppressors being in place (The Register: the strike 'overloaded suppressors on the mechanical cooling system, shutting it down'). The surge suppression was present but its margin was inadequate to ride through a lightning-scale swell without cooling loss. Note the asymmetry: the electrical protection worked for IT power (transfer to generator) but the mechanical/cooling plant tripped — a redundancy that did not take over. Thermal layer — buffer depletion and threshold breach. The cooling system incorporated a load-dependent thermal buffer — engineered thermal mass/inertia that let the hall maintain operational temperatures for a finite window after active cooling stopped. This buffer is load-dependent: the higher the IT heat load, the faster it depletes. Once depleted, datacenter temperature exceeded safe operational thresholds. The RCA does not disclose the actual air temperatures reached or the numeric thresholds, so any specific degree figure would be unsourced; what is documented is that the breach was detected by environmental/thermal monitoring and triggered an automated device shutdown. Protection layer — automated integrity shutdown, partially defeated. The safety mechanism here is not fire suppression but an automated shutdown of devices intended to preserve infrastructure and data integrity. It activated. But temperatures increased so quickly in parts of the datacenter that some hardware was damaged before it could shut down. The damage tally: a significant number of storage servers were damaged, as well as a small number of network devices and power units. First customer-visible effect was at 09:29 UTC as storage servers began shutting down on unsafe temperatures. Onsite teams took actions to prevent further damage — including transferring the rest of the datacenter to generators, stabilizing the power supply — and did so while storms were still active; no evacuation and no emergency-services dispatch appears in the record. Logical/service layer — where the real spread happened. Physical damage stayed confined to hardware in the affected zones of one datacenter, but the outage propagated far beyond via single-region dependencies. ASM keeps resource metadata in multiple locations but its primary site is South Central US and it does not support automatic failover — this led to the widest impact for customers outside the region, not mitigated until the associated South Central US storage servers were brought back online at 01:10 UTC on 5 September. A North America Azure AD site sat in the affected datacenter; as infrastructure shut down, authentication traffic began automatically routing to other sites (AAD self-healed, mitigated ~14:40 UTC on 4 September per the RCA). Separately, the Azure status page itself buckled: increased traffic combined with incorrect auto-scale configuration prevented the web app from scaling, causing intermittent 500 errors until ~23:05 UTC. Recovery engineering. Microsoft chose recovery over failover deliberately: a failover would have resulted in limited data loss due to the asynchronous nature of geo-replication, so the decision was made to work towards recovery of data rather than fail over. Recovery was staged: first restore the Azure Software Load Balancers (SLBs) for the storage scale units, then recover the storage servers and their data — replacing failed infrastructure components, migrating customer data from damaged servers to healthy servers, and validating that no recovered data was corrupted.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home