← All incidents
Incident dossier · Rank #39

London Heatwave Cooling Failures — Google Cloud & Oracle (July 2022)

Google Cloud & Oracle 2022-07-19 24h 0m core impact Cooling

During a record UK heatwave with temperatures above 40°C, cooling systems failed at Google Cloud (europe-west2) and Oracle (UK South) London data centers, forcing hardware/VM shutdowns and regional service loss — an early climate-driven thermal event.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Cooling (2022-07-19)Trigger · Cooling2022-07-192022-07-19Primary fault at Google Cloud & Oracle — Google Cloud europe-west2 & Oracle UK South (London)Google & OracleGoogle Cloud europe-west2 & Oracle UK South (London)Google Cloud europe-west2 & OracleDownstream service degraded by the fault: Google Cloud europe-west2 (partial)Google Cloud europe-west2Downstream service degraded by the fault: Oracle Cloud UK South (partial)Oracle Cloud UK South

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Google Cloud & Oracle
Data center
Google Cloud europe-west2 & Oracle UK South (London)
Location
London, United Kingdom
Date
2022-07-19

Impact & scale

Users affected
Google Cloud europe-west2 and Oracle UK South customers
Financial
Not published
Scope
Climate-driven cooling loss (regional)
Services / systems down
  • Google Cloud europe-west2 (partial)
  • Oracle Cloud UK South (partial)

Impact data & metrics

Record ambient outside-air temperature (UK national record)40.3C (104.5F) - highest ever recorded in the UK
Google zone affectedeurope-west2-a (europe-west2-b and -c not impacted for VMs)
Google VMs terminated in europe-west2-a'a small set' (public status page); ~35% of the zone's VMs per The Register's report of Google's fuller incident report
Google time to 'resolved for all affected users'~14h10m (06:33 -> 20:43 US/Pacific, 2022-07-19)
Google official incident window (US/Pacific)Began 2022-07-19 06:33, ended 2022-07-20 21:20 (~38h47m to full closure, storage tail)
Oracle site affectedUK South (London) - subset of cooling infrastructure
Fire-suppression agent discharged (gaseous/water-mist/sprinkler)0 - none activated (no fire occurred)
Personnel evacuated0 - no evacuation reported (equipment-protection shutdown, not life-safety)
Emergency-services / fire-brigade calls and arrivals0 - none reported
Lingering storage damage (Google)Small percentage of replicated Persistent Disks in single-redundant mode; small number of HDD-backed PD volumes exhibiting IO errors
Operators impacted by the same heat event (same day)2 - Google (europe-west2-a) and Oracle (UK South)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 5Users affected (0–10) — breadth of the user/customer population impacted. — scored 5/10.Financial 4Financial impact (0–10) — direct + consequential cost. — scored 4/10.Duration 5Outage duration (0–10) — how long service was degraded/down. — scored 5/10.Blast 6Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 6/10.
Magnitude 5.2 = blast 6×0.35 + users 5×0.25 + financial 4×0.20 + duration 5×0.20 (sub-scores 0–10 · weighted composite)

Record >40C UK heat overwhelmed cooling, forcing shutdowns at Google Cloud and Oracle London DCs — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events2022-07-19 (daytime, UK) · TRIGGER — UK ambient air reaches an all-time national record of 40.3C (104.5F); London experiences its first-ever day above 40C, driving cooling units toward their design limits.TRIGGER2022-07-19 (dayt2022-07-19 06:33 US/Pacific (13:33 UTC) · TRIGGER — Google incident begins: a cooling-related failure in one building hosting zone europe-west2-a; The Register (2022-08-01) reported this as a simultaneous failure of multiple redundant cooling systems under record ambient.TRIGGER2022-07-19 06:33 US/Pacific (13:33 U2022-07-19 (post-06:33 PDT) · DETECTION — Data-centre environmental/BMS temperature monitoring (NOT fire/smoke detection - no combustion) registers loss of cooling; the building could not maintain a safe operating temperature. No fire alarm - a thermal, not fire, condition.DETECTION2022-07-19 (post-06:33 PD2022-07-19 (immediately after detection) · MITIGATION — Protective thermal load-shed (de-energisation analogue): Google powers down part of the zone and limits preemptible launches to prevent machine damage. NO fire suppression discharged - no gaseous/water-mist/sprinkler activation, because there was no fire.MITIGATION2022-07-19 (imme2022-07-19 · IMPACT — Compute Engine terminates VMs in the impacted datacenter - a small set per the public status page; ~35% of europe-west2-a VMs per The Register's report of Google's fuller incident report.IMPACT2022-07-192022-07-19 · CASCADE — Recovery-procedure error widens blast radius: engineers inadvertently modify traffic routing for internal services to avoid all three europe-west2 zones rather than just the impacted europe-west2-a.CASCADE2022-07-192022-07-19 ~16:15 UTC (~09:15 US/Pacific) · DETECTION — Google publicly acknowledges the fault: 'There has been a cooling-related failure in one of our buildings that hosts zone europe-west2-a for region europe-west2.'DETECTION2022-07-19 ~16:15 U2022-07-19 (UK afternoon) · TRIGGER — Oracle UK South (London): a subset of cooling infrastructure experiences an issue as a result of unseasonal temperatures - a second, independent London campus hit by the same heat event.TRIGGER2022-07-19 (UK a2022-07-19 16:38 UTC · DETECTION — Oracle acknowledges the cooling problem at its UK South (London) data centre.DETECTION2022-07-19 16:38 U2022-07-19 (after acknowledgement) · MITIGATION — Oracle protective power-down: identifying 'service infrastructure that can be safely powered down to prevent additional hardware failures' - thermal load-shed, not fire response.MITIGATION2022-07-19 (afte2022-07-19 (throughout) · IMPACT — NEGATIVE FORENSICS - no personnel evacuation reported (equipment-protection shutdown, not a life-safety event) and no emergency-services / fire-brigade call or arrival is reported; response was operator engineering teams restoring cooling and rebuilding capacity.IMPACT2022-07-19 (thro2022-07-19 (afternoon, per The Register) · RECOVERY — Google cooling reported restored (~14:13 US/Pacific per The Register 2022-08-01) and capacity rebuild begins in europe-west2-a.RECOVERY2022-07-19 (afternoon, per 2022-07-19 20:43 US/Pacific (03:43 UTC 20 Jul) · RESTORED — Google: 'the issue has been resolved for all affected users as of Tuesday, 2022-07-19 20:43 US/Pacific' (~14h10m from 06:33 start), though some storage issues persist requiring support contact.RESTORED2022-07-19 20:43 US/Pacific (03:43 U2022-07-20 (recovery update) · IMPACT — Storage long-tail: a small number of HDD-backed Persistent Disk volumes still exhibit IO errors, with a small percentage of replicated PDs running in single-redundant mode - genuine media damage from the thermal event/abrupt power-down.IMPACT2022-07-20 (reco2022-07-20 ~03:57 UTC (per The Register follow-up) · RECOVERY — Oracle data-centre cooling infrastructure reported restored and temperatures returned to normal operating levels.RECOVERY2022-07-20 ~03:57 U2022-07-20 21:20 US/Pacific · RESTORED — Google official incident window closes. Status-page window: began 2022-07-19 06:33, ended 2022-07-20 21:20 US/Pacific (~38h47m to full closure); the extra tail beyond the 20:43 Jul 19 'resolved for all users' mark was storage remediation.RESTORED2022-07-20 21:20

Root cause

SPECIFIC MECHANISM (Google, europe-west2-a, London): The proximate cause was a cooling-capacity collapse in one building hosting zone europe-west2-a on the UK's hottest day on record. Google's PUBLIC incident report (status.cloud.google.com, incident XVq5om2XEDSqLtJZUvcH) states there was "a cooling related failure in one of our buildings that hosts zone europe-west2-a for region europe-west2." The Register's 2022-08-01 reporting of Google's fuller (customer-facing) incident report characterised the mechanism more sharply — a simultaneous failure of multiple, redundant cooling systems combined with the extraordinarily high outside temperature such that the building "could not maintain a safe operating temperature." NOTE (verification): this stronger "simultaneous / multiple redundant systems" wording appears in The Register's 2022-08-01 write-up and NOT in the public status page (which says only "a cooling related failure"); it is retained as attributed reporting, not re-verified verbatim in this pass. When cooling capacity collapsed, Compute Engine executed a protective thermal load-shed and Google "powered down part of the zone" to prevent machine damage (The Register, 2022-07-19). The Register (2022-08-01) reported ~35% of europe-west2-a's VMs were terminated; the public status page confirms only "abnormal VM terminations for a small set." LATENT ROOT — DESIGN / REDUNDANCY MARGIN: The pairing "simultaneous" + "redundant" is the tell of a latent design defect — genuine N+1/N+2 redundancy should not permit two backup-capable units to fail at effectively the same moment, which points to a common-mode driver (shared record ambient) or a single point of failure. Trade analysis (The Register, 2022-08-22, citing Tech Monitor) framed the units as failing when "required to operate above their design limits," i.e. a climate-resilience / design-margin gap: the plant lacked head-room for a 40°C day. The UK reached an all-time record 40.3°C (104.5°F) on 2022-07-19 (The Register, 2022-07-19). CAVEAT: the "two cooler units above design limits," "single point of failure," and Equinix worst-case-factory-test contrast originate in trade reporting (Tech Monitor / The Register Aug 22) that could not be independently re-fetched this pass; they are inference/analysis, not operator-disclosed fact. MAINTENANCE / INSPECTION / TESTING LAPSE: NONE EVIDENCED in any public source. Neither Google's report nor Oracle's status notes attribute the failure to deferred maintenance, a botched service, an untested standby chiller, or a work-procedure error. The disclosed cause is environmental (record ambient) plus a cooling-system failure. Disclosed-uncertainty point: the record is silent on whether the redundant units had ever been function-tested under near-design ambient, and equipment specifics (chiller/CRAC/CRAH make, model, age, refrigerant chemistry; whether free-cooling economisers or mechanical DX/chilled-water plant failed) were NOT disclosed by either operator for either site. ORACLE (UK South, London): Same climate trigger, less detail. "As a result of unseasonal temperatures in the region, a subset of cooling infrastructure within the UK South (London) Data Centre has experienced an issue" (Oracle status, via The Register 2022-07-19), forcing a protective power-down of "service infrastructure that can be safely powered down to prevent additional hardware failures." Oracle named no equipment, age, refrigerant, or maintenance history.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What actually happened - and what did not

On 2022-07-19, the UK's hottest day on record (40.3C), cooling failed at two independent London cloud campuses - Google's europe-west2-a and Oracle's UK South. This was a thermal cooling-capacity collapse, not a fire. There was no ignition, no combustion, no smoke/heat detection, no suppression discharge, no evacuation, and no emergency-services response reported by any source. The correct analogues are environmental/BMS temperature monitoring in place of fire detection, and deliberate thermal load-shedding / de-energisation in place of fire suppression. Google confirmed 'a cooling related failure in one of our buildings that hosts zone europe-west2-a', and 'powered down part of the zone' to prevent machine damage.

The latent root: design margin versus a changing climate

The units were, per trade reporting, 'required to operate above their design limits' - a design-margin gap rather than a maintenance failure. The plant's worst-case ambient envelope appears to have been set below the emerging UK climate reality; London had never before exceeded 40C. This reframes the event from an equipment fault to a climate-resilience shortfall: cooling designed and factory-tested for historical worst-case conditions is now under-provisioned. (The 'above design limits' framing is trade analysis via Tech Monitor/The Register and was not operator-confirmed or re-verified this pass.)

Redundancy that did not deliver

Google's fuller report (via The Register, 2022-08-01) described a simultaneous failure of multiple redundant cooling systems. Simultaneity among 'redundant' units is the tell of a common-mode driver - here the shared record ambient - or a single point of failure. Genuine N+1/N+2 redundancy should not permit backup-capable units to fail together; that they did indicates the redundancy was duplicated rather than diverse and independent. This is the single most important engineering lesson of the event.

Self-inflicted blast radius: the failover-tooling error

The physical impact was confined to one building/hall of europe-west2-a (europe-west2-b and -c were explicitly unaffected for VMs). The damaging secondary spread was operational: engineers inadvertently rerouted internal-service traffic away from all three europe-west2 zones rather than just the impacted one, widening customer impact. This procedural error - not the cooling failure - drove much of the customer-visible blast radius, underscoring recovery-tooling discipline. (Reported by The Register; not on the public status page.)

Storage long-tail and the cost of emergency shutdown

Protective de-energisation is not free. Google's status page confirms 'a small percentage of replicated Persistent Disk devices are now running in single redundant mode' and 'a small number of HDD backed Persistent Disk volumes are still experiencing impact and will exhibit IO errors.' The abrupt power-down inflicted genuine storage-media damage and left durability degraded, extending the official incident window (06:33 Jul 19 to 21:20 Jul 20 US/Pacific) well past the 'resolved for all users' mark at 20:43 Jul 19.

Two operators, one heat event, and the disclosure gap

Both Google and Oracle cooling plant failed the same day under the same heat - a sector-level design-assumption signal, not a one-off. Yet disclosure was thin: neither operator released equipment make/model/age, refrigerant chemistry, plant type (free-cooling economiser vs mechanical DX/chilled-water), maintenance/test history, MW lost, rack counts, or cost. No maintenance or inspection lapse is evidenced anywhere; equally, whether the redundant units had ever been function-tested under near-design ambient is unknown. These remain disclosed-uncertainty points, not findings.

Technical deep-dive

This was a thermal-overload / cooling-capacity-collapse event at two independent London cloud campuses on the UK's hottest day on record — NOT a fire. Every fire-forensic axis (ignition, detection-of-combustion, suppression discharge, evacuation, emergency-services response) is therefore a documented NEGATIVE, and the correct analogues are environmental/BMS temperature monitoring in place of smoke/heat detection, and deliberate thermal load-shedding / de-energisation in place of fire suppression. FAILURE CHAIN (Google europe-west2-a): (1) Ambient forcing — record 40.3°C outside air (The Register, 2022-07-19) pushed mechanical cooling toward/above rated design limits. (2) Cooling failure — a cooling-related failure in the building hosting europe-west2-a (public status page); The Register (2022-08-01) reported this as a simultaneous failure of multiple redundant cooling systems, implying a common-mode driver or single point of failure so that intended redundancy did not deliver. (3) Sensed condition — data-centre environmental/BMS temperature monitoring registered rising hall temperature (NOT fire/smoke detection; no combustion). (4) Protective action — Google "powered down part of the zone" and limited preemptible launches "to prevent machine damage" (status page / The Register). This is thermal load-shedding to prevent uncontrolled hardware/thermal damage — the physical analogue of de-energisation, NOT fire suppression. No gaseous, water-mist, or sprinkler system was discharged, because there was no fire. BLAST-RADIUS AMPLIFICATION (reported operational error): The physical "spread" was purely thermal and confined to the affected building/hall of europe-west2-a (europe-west2-b and -c were explicitly not impacted for VMs — status page). The Register (2022-08-01) reported the damaging SECONDARY spread was operational: engineers inadvertently modified traffic routing for internal services to avoid all three europe-west2 zones rather than just europe-west2-a, widening customer impact. This traffic-routing detail is not on the public status page and is retained as attributed reporting; it is a procedural (not maintenance) fault and a key recovery-tooling lesson. RECOVERY & LONG-TAIL DAMAGE: Google confirmed "the issue has been resolved for all affected users as of Tuesday, 2022-07-19 20:43 US/Pacific," ~14h10m after the 06:33 US/Pacific start; the official incident window on the status page runs 2022-07-19 06:33 to 2022-07-20 21:20 US/Pacific (~38h47m to full closure), the extra tail driven by storage. NOTE: The Register's precise "18h23m primary / 35h15m long-tail" figures could not be verified and do not reconcile with the public window, so status-page-derived figures are used instead. Genuine media damage is confirmed on the status page: "A small percentage of replicated Persistent Disk devices are now running in single redundant mode," and later "a small number of HDD backed Persistent Disk volumes are still experiencing impact and will exhibit IO errors" — attributable to the thermal event and abrupt power-down. Oracle reported (per The Register follow-up, not re-verified this pass) cooling restored and temperatures normal by ~03:57 UTC 20 Jul, with a subset of Oracle Integration Cloud still impacted afterward. EVIDENCE GAPS (disclosed uncertainty): No equipment make/model/age/refrigerant chemistry was released for either operator; neither stated whether economiser/free-cooling or mechanical DX/chilled-water plant failed; no maintenance/test history was published; no MW-lost, rack-count, coolant-flow, or dollar-cost figures were disclosed. No fire, no evacuation, no emergency-services response was reported by any source.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home