← Home
Root-only · Post-incident dossier

Data-Center Incidents — Case Library

A structured, source-cited library of major data-center and cloud-infrastructure incidents worldwide — ranked by magnitude, each with a full sequence of events, root-cause analysis, correction-of-errors (COE), lessons learnt and engineering improvements. Every fact is traced to a public post-incident report.

43 catalogued·30 with official RCA·provenance-mandatory·root access
About this library · how to read it

What this is

A root-gated, source-cited engineering reference on major data-center and cloud-infrastructure incidents worldwide — each a structured post-incident dossier (sequence of events, root-cause analysis, correction-of-errors, lessons learnt, engineering improvements). It exists for internal engineering education; every material fact is traced to a public post-incident report, and where an authoritative cause was never published the dossier says so rather than invent one.

How incidents are ranked

By a transparent magnitude composite — a weighted sum of four sourced sub-scores (each 0–10): blast radius 0.35 · users 0.25 · financial 0.20 · duration 0.20. Hover any magnitude bar or a dossier's radar for the exact breakdown. Ranking never uses a hidden score.

What “official RCA” means

The official-RCA filter and the count above mark incidents backed by a genuine vendor post-incident review, regulator/government report, or court record — not press reporting. The provenance gate refuses to set that flag unless such a source is actually cited.

Reading the visualizations

  • Where these happened — a real map at each incident's origin site; gold = facility failures (power/cooling/fire), green = network/logical. Co-located incidents are fanned onto a small ring; hover a marker for the brief.
  • Ranked bars — magnitude, coloured by severity (red = most severe).
  • Risk map — blast radius (x) × outage duration (y); top-right is the worst quadrant. Dots sit at their true scores; hover any dot for its values.
  • Semantic map — a projection of the research vector-index: incidents near each other share a failure signature.
  • Per dossier — a failure-cascade block diagram, a magnitude radar, and a phased SOE timeline; hover elements for detail, and hover underlined terms for definitions.

Access

Root-only. The library is excluded from the public sitemap and search index.

Incident Intelligence

Global data-center incident landscape · risk map & semantic analysis

All-time · 2011–2026
Incidents catalogued
0
deep dossiers · span 2011–2026
Critical (magnitude ≥ 8.5)
0
11 more rated High
Countries affected
0
distinct national jurisdictions
Avg blast radius
0/ 10
weighted spread of impact
Avg outage duration
116h 6m
core-impact window (mean)
Risk map · blast radius × severity
Facility · power / cooling / fireNetwork / logical · software / network / humanbigger dot = higher magnitude · tap a dot to open the incident
Failure signatures
  • Fire / battery0mag 7.1
  • Capacity / scaling / quota0mag 7.3
  • Power / electrical0mag 6.8
  • Software / config / deploy0mag 7.4
  • Human error / procedure0mag 7.3
  • Network / BGP / DNS0mag 7.3
  • Cooling / thermal0mag 6.5
Duration distribution
  • < 15m0%
  • 15–60m2%
  • 1–6h21%
  • 6–24h47%
  • 1–3d16%
  • > 3d14%
Source public post-incident reportsCoverage 2011–2026Catalogued 43 incidentsOfficial RCA 30 of 43Methodology →

All incidents

Incident IDTitleSeverityStatusData center / regionStartDurationBlastImpactTags
INC-2024-001CrowdStrike Falcon Channel File 291: The Global Windows BSOD Outage of July 2024
On 19 July 2024 CrowdStrike pushed a Rapid Response Content update to Channel File 291 t…
CriticalOfficial RCA🌐Global Windows fleet (all clouGlobal2024-07-191h 18m10.0/10
SoftwareHuman error
INC-2022-002SK C&C Pangyo Data Center Fire and the National Kakao/Naver Outage
On 15 October 2022, a lithium-ion UPS battery on the third basement level (B3) of the SK…
CriticalOfficial RCA🇰🇷SK C&C Pangyo Data CenterSeongnam (Pangyo), South Korea2022-10-15127h 30m10.0/10
FirePower
INC-2025-003NIRS Daejeon Government Data Center: UPS Lithium-Ion Battery Fire During Relocation Work
On the evening of 26 September 2025, a lithium-ion UPS battery that had been disconnecte…
CriticalOfficial RCA🇰🇷NIRS Daejeon National InformatDaejeon, South Korea2025-09-2610h 0m10.0/10
FirePower
INC-2021-004Meta Global BGP + DNS Outage — Backbone Self-Disconnection Takes Facebook, Instagram & WhatsApp Offline Worldwide
On 4 October 2021 a single command issued during routine backbone maintenance — intended…
CriticalOfficial RCA🇺🇸Facebook Global Backbone / AutMenlo Park, USA2021-10-047h 11m10.0/10
NetworkHuman error
INC-2025-005Google Cloud Global Service Control / IAM Outage — June 12, 2025
A latent, unguarded code path in Google Cloud's Service Control quota system — added on …
CriticalOfficial RCA🇺🇸Google Cloud — Service ControlUSA2025-06-127h 27m10.0/10
Software
INC-2025-006AWS US-EAST-1 DynamoDB DNS Race Condition Cascades Across the Internet
A latent race condition in DynamoDB's automated DNS management system left two DNS Enact…
CriticalOfficial RCA🇺🇸US-EAST-1 (Northern Virginia)Ashburn, United States2025-10-2014h 32m10.0/10
SoftwareNetwork
INC-2021-007AWS US-EAST-1 Internal Network Congestion Outage (December 7, 2021)
On December 7, 2021, an automated capacity-scaling operation on a service hosted in AWS'…
CriticalOfficial RCA🇺🇸AWS US-EAST-1 (Northern VirginAshburn, USA2021-12-076h 52m9.0/10
NetworkSoftware
INC-2021-008OVHcloud SBG2 Strasbourg Data-Center Fire (10 March 2021)
In the early hours of 10 March 2021 a fire broke out in a ground-floor energy room at OV…
CriticalOfficial RCA🇫🇷SBG2 (Strasbourg campus)Strasbourg, France2021-03-109h 27m9.0/10
FirePower
INC-2017-009AWS S3 US-EAST-1 Service Disruption
An authorized S3 engineer running an established capacity-removal playbook mistyped one …
CriticalOfficial RCA🇺🇸US-EAST-1 region (Northern VirAshburn (approx.), United States2017-02-284h 35m10.0/10
Human errorSoftware
INC-2016-010Dyn Managed DNS Outage — Mirai IoT Botnet DDoS (October 21, 2016)
On Friday 21 October 2016, Dyn's managed (authoritative) DNS platform was hit by a large…
CriticalOfficial RCA🇺🇸Dyn Managed DNS Platform (authManchester, USA2016-10-2111h 0m10.0/10
NetworkSoftware
INC-2025-011Azure Front Door Global Outage: Inadvertent Cross-Build Config Change Trips Latent Data-Plane Crash Bug (29 Oct 2025)
On 29 October 2025 an inadvertent customer configuration change to Azure Front Door (AFD…
HighOfficial RCA🌐Azure Front Door global edge fGlobal2025-10-298h 24m10.0/10
SoftwareHuman error
INC-2016-012Delta Air Lines Atlanta Data Center Power Failure Grounds the Global Fleet
In the early hours of 8 August 2016, a routine scheduled transfer of load onto a backup …
HighPress-sourced🇺🇸Delta Technology Data Center, Atlanta, United States2016-08-086h 0m9.0/10
PowerFire
INC-2022-013Rogers Communications Canada Nationwide Network Outage (July 2022)
A maintenance change to Rogers' core IP network removed a routing filter, flooding the r…
HighOfficial RCA🇨🇦Rogers national IP core networToronto, Canada2022-07-0815h 0m9.0/10
NetworkSoftwareHuman error
INC-2020-014AWS Kinesis OS Thread-Limit Cascade — US-EAST-1 (Nov 25, 2020)
A small addition of capacity to the Kinesis Data Streams front-end fleet in US-EAST-1 pu…
HighOfficial RCA🇺🇸AWS US-EAST-1 (Kinesis Data StAshburn, USA2020-11-2517h 8m9.0/10
Software
INC-2024-015Google Cloud Accidentally Deletes UniSuper's Entire GCVE Private Cloud
During the initial provisioning of UniSuper's Google Cloud VMware Engine (GCVE) Private …
HighOfficial RCA🇦🇺Google Cloud VMware Engine (GCAustralia2024-05-02216h 0m10.0/10
Human errorSoftware
INC-2025-016Cloudflare Global Outage: Oversized Bot Management Feature File (November 18, 2025)
On 18 November 2025 an 11:05 UTC ClickHouse database permissions change exposed r0-datab…
HighOfficial RCA🌐Cloudflare global edge networkGlobal2025-11-185h 46m9.0/10
Software
INC-2020-017Google Global Authentication Outage — User ID Service Quota Collapse (Dec 14, 2020)
On December 14, 2020, an automated quota-management system wrongly cut capacity on the a…
HighOfficial RCA🇺🇸Google global identity (User IUSA2020-12-1447m10.0/10
Software
INC-2020-018Equinix LD8 Docklands UPS Failure: ~17.5-Hour Power Outage Disrupts LINX and ~150 Members
On 18 August 2020 a faulty UPS system at Equinix's LD8 Docklands data centre — home to t…
HighOfficial RCA🇬🇧LD8 (Docklands)London, United Kingdom2020-08-1817h 30m9.0/10
Power
INC-2026-019NorthC Almere Data Center Fire: A Compartment Blaze That Took Out the Facility Power Plant and Rippled Across Dutch Public Services
On 7 May 2026 (around 08:45 local time) a fire broke out in a rear technical compartment…
HighPress-sourced🇳🇱NorthC Almere (Rondebeltweg)Almere, Netherlands2026-05-07144h 0m9.0/10
FirePower
INC-2020-020Azure Active Directory Global Authentication Outage (September 28, 2020)
On 28 September 2020 a service update meant only for an internal Azure AD validation tes…
HighOfficial RCA🇺🇸Azure Active Directory — globaUSA2020-09-285h 0m9.0/10
Software
INC-2023-021Microsoft Azure / Microsoft 365 Global WAN Outage — IGP-Purge Command Cascade (Jan 25 2023)
A single IGP-database-purge command, added to a WAN change procedure that was never re-t…
HighOfficial RCA🌐Global WAN backbone (Madrid caMadrid, Spain2023-01-255h 35m9.0/10
Networkconfig-errorchange-management
INC-2011-022AWS EBS Re-Mirroring Storm and Stuck Volumes in US-East-1
A routine network-capacity upgrade in one US-East-1 Availability Zone was executed incor…
MediumOfficial RCA🇺🇸US-East-1 (Northern Virginia)Ashburn area, Northern Virginia, USA2011-04-2190h 43m7.0/10
Networkconfig-errorSoftwarestorage
INC-2023-023Optus Australia Nationwide Outage — Routing Change Cascade (November 8, 2023)
Following a routine Singtel-network software upgrade, a large, unexpected flood of routi…
MediumPress-sourced🇦🇺Optus national network coreSydney, Australia2023-11-0814h 0m8.0/10
NetworkSoftware
INC-2023-024Alibaba Cloud Global All-Region Outage — Auth Control-Plane Failure
On 12 November 2023 Alibaba Cloud's shared global authentication/authorization (Auth) co…
MediumPress-sourced🇨🇳Global control plane (Auth/IAMHangzhou, China2023-11-123h 27m9.0/10
Softwareconfig-errorcontrol-planeauthentication
INC-2023-025Cloudflare Control-Plane & Analytics Outage — Flexential PDX-04 Power Failure (Nov 2023)
On 2 November 2023 a Portland General Electric (PGE) unplanned maintenance event dropped…
MediumOfficial RCA🇺🇸Flexential PDX-04Hillsboro, United States2023-11-0240h 41m8.0/10
PowerHuman error
INC-2018-026Azure South Central US: Lightning-Induced Cooling Loss and Automated Integrity Shutdown (September 2018)
On 4 September 2018, a high-energy thunderstorm with lightning struck southern Texas nea…
MediumOfficial RCA🇺🇸South Central US region (San ASan Antonio, United States2018-09-0438h 20m7.0/10
CoolingPower
INC-2024-027Red Sea Subsea Cable Cuts — Rubymar Anchor Severs AAE-1, Seacom/TGN-EA and EIG off Yemen
On 24 February 2024 three submarine cable systems — Seacom (which Kentik groups with TGN…
MediumPress-sourced🇾🇪Red Sea cable corridor (Bab-elBab-el-Mandeb, Yemen2024-02-243408h 0m8.0/10
NetworkSupply chain
INC-2024-028AT&T Nationwide Wireless Outage — Network Misconfiguration (February 22, 2024)
During a network expansion, an AT&T engineer used an incorrect process that introduced a…
MediumOfficial RCA🇺🇸AT&T national wireless core neDallas, Texas, United States2024-02-2212h 0m8.0/10
NetworkSoftwareHuman error
INC-2025-029CyrusOne Aurora Cooling Failure Halts CME Group Trading (November 2025)
A human-error cooling-tower switchover during a cold-weather changeover triggered a chil…
MediumPress-sourced🇺🇸CyrusOne Aurora (Chicago-area)Aurora, Illinois, United States2025-11-0111h 0m7.0/10
CoolingHuman error
INC-2026-030AWS US-EAST-1 Cooling/Thermal Event and use1-az4 Power Loss
On 7 May 2026 cooling systems failed in AWS's use1-az4 availability zone in US-EAST-1 (N…
MediumPress-sourced🇺🇸US-EAST-1 (use1-az4), NorthernAshburn, United States2026-05-0720h 25m6.0/10
CoolingPower
INC-2017-031British Airways Boadicea House Power Surge: A 15-Minute Fault That Grounded Heathrow and Gatwick for Days
On the morning of Saturday 27 May 2017, an engineer working inside British Airways' Boad…
MediumPress-sourced🇬🇧Boadicea House (BoHo) data cenLondon, United Kingdom2017-05-2748h 0m7.0/10
PowerHuman error
INC-2022-032Cloudflare June 2022 Outage — Config Change Withdraws 19 Core Data Centers
A network-configuration change during a resilience upgrade to Cloudflare's Multi-Colo Po…
MediumOfficial RCA🌐Cloudflare 19 busiest core datGlobal, Global2022-06-211h 15m8.0/10
NetworkSoftware
INC-2026-033Google Cloud India Disruption After STT GDC Delhi Data-Center Fire
On or about 5 June 2026, a fire at a third-party data-center facility in New Delhi force…
MediumOfficial RCA🇮🇳Third-party colocation facilitNew Delhi, India2026-06-05477h 53m6.0/10
FirePowerNetwork
INC-2021-034Fastly Global CDN Outage — Latent Software Bug (June 8, 2021)
An undiscovered software bug introduced in a mid-May Fastly software deployment lay dorm…
MediumOfficial RCA🌐Fastly global edge CDN networkGlobal, Global2021-06-081h 0m8.0/10
SoftwareNetwork
INC-2025-035PlayStation Network ~24-Hour Global Outage (February 2025)
An approximately 24-hour global PlayStation Network outage on 7–8 February 2025 locked m…
LowPress-sourced🇺🇸PlayStation Network backend seSan Mateo, California, United States2025-02-0724h 0m6.0/10
NetworkSoftware
INC-2026-036Azure West US 2 Grid-Disturbance Cooling Loss and Thermal Shutdown (May 2026)
A grid power-quality disturbance reduced cooling capacity across multiple Azure West US …
LowOfficial RCA🇺🇸Azure West US 2 (Quincy, WashiQuincy, Washington, United States2026-05-018h 0m6.0/10
PowerCooling
INC-2022-037Virginia 'Data Center Alley' Heat-Dome Cooling Stress (July 2022)
A July 2022 heat dome pushed dozens of Northern Virginia 'Data Center Alley' facilities …
LowPress-sourced🇺🇸Northern Virginia data-center Ashburn (Loudoun), Virginia, United States2022-07-246h 0m6.0/10
Cooling
INC-2025-038X (Twitter) Oregon Data-Center Power-Cabinet Fire (May 2025)
A fire that began inside a power cabinet at a Digital Realty Oregon facility hosting X (…
LowOfficial RCA🇺🇸Digital Realty Oregon facilityHillsboro, Oregon, United States2025-05-2448h 0m6.0/10
FirePower
INC-2022-039London Heatwave Cooling Failures — Google Cloud & Oracle (July 2022)
During a record UK heatwave with temperatures above 40°C, cooling systems failed at Goog…
LowOfficial RCA🇬🇧Google Cloud europe-west2 & OrLondon, United Kingdom2022-07-1924h 0m6.0/10
Cooling
INC-2024-040Azure China North 3 ~50-Hour Regional Outage (2024)
A power/cooling failure in Azure China North 3 (operated by 21Vianet) caused a roughly 5…
LowPress-sourced🇨🇳Azure China North 3 region (HeHebei, China2024-11-0150h 0m4.0/10
PowerCooling
INC-2024-041Digital Realty Singapore Lithium-Ion UPS Fire (2024)
A lithium-ion UPS battery fire at a Singapore data center forced an emergency power shut…
LowPress-sourced🇸🇬Digital Realty Singapore data Singapore, Singapore2024-09-0110h 0m5.0/10
FirePower
INC-2022-042Google Council Bluffs (Iowa) Electrical Explosion / Arc-Flash (August 2022)
An electrical explosion (arc-flash) at Google's Council Bluffs, Iowa data center injured…
LowOfficial RCA🇺🇸Google Council Bluffs data cenCouncil Bluffs, Iowa, United States2022-08-083h 0m5.0/10
PowerFire
INC-2023-043Digital Realty Los Angeles (El Segundo) Data-Center Fire (May 2023)
A fire in server/equipment space at Digital Realty's Los Angeles (El Segundo) data cente…
LowPress-sourced🇺🇸Digital Realty LAX (El SegundoEl Segundo, California, United States2023-05-2112h 0m4.0/10
Fire

Ranking is a transparent composite of sourced sub-scores (blast radius 35% · users 25% · financial 20% · duration 20%). Summaries are original and substantially shorter than their sources; short attributed excerpts on each incident page are provenance only. This library is for engineering education and does not reproduce source material in full.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home