London Heatwave Cooling Failures — Google Cloud & Oracle (July 2022)
During a record UK heatwave with temperatures above 40°C, cooling systems failed at Google Cloud (europe-west2) and Oracle (UK South) London data centers, forcing hardware/VM shutdowns and regional service loss — an early climate-driven thermal event.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Google Cloud & Oracle
- Data center
- Google Cloud europe-west2 & Oracle UK South (London)
- Location
- London, United Kingdom
- Date
- 2022-07-19
Impact & scale
- Users affected
- Google Cloud europe-west2 and Oracle UK South customers
- Financial
- Not published
- Scope
- Climate-driven cooling loss (regional)
- Google Cloud europe-west2 (partial)
- Oracle Cloud UK South (partial)
Impact data & metrics
| Record ambient outside-air temperature (UK national record) | 40.3C (104.5F) - highest ever recorded in the UK |
| Google zone affected | europe-west2-a (europe-west2-b and -c not impacted for VMs) |
| Google VMs terminated in europe-west2-a | 'a small set' (public status page); ~35% of the zone's VMs per The Register's report of Google's fuller incident report |
| Google time to 'resolved for all affected users' | ~14h10m (06:33 -> 20:43 US/Pacific, 2022-07-19) |
| Google official incident window (US/Pacific) | Began 2022-07-19 06:33, ended 2022-07-20 21:20 (~38h47m to full closure, storage tail) |
| Oracle site affected | UK South (London) - subset of cooling infrastructure |
| Fire-suppression agent discharged (gaseous/water-mist/sprinkler) | 0 - none activated (no fire occurred) |
| Personnel evacuated | 0 - no evacuation reported (equipment-protection shutdown, not life-safety) |
| Emergency-services / fire-brigade calls and arrivals | 0 - none reported |
| Lingering storage damage (Google) | Small percentage of replicated Persistent Disks in single-redundant mode; small number of HDD-backed PD volumes exhibiting IO errors |
| Operators impacted by the same heat event (same day) | 2 - Google (europe-west2-a) and Oracle (UK South) |
Magnitude profile
Record >40C UK heat overwhelmed cooling, forcing shutdowns at Google Cloud and Oracle London DCs — sub-scores ESTIMATED from public impact reporting, pending deep research.
Sequence of events (SOE)
- TRIGGER UK ambient air reaches an all-time national record of 40.3C (104.5F); London experiences its first-ever day above 40C, driving cooling units toward their design limits.
- TRIGGER Google incident begins: a cooling-related failure in one building hosting zone europe-west2-a; The Register (2022-08-01) reported this as a simultaneous failure of multiple redundant cooling systems under record ambient.
- DETECTION Data-centre environmental/BMS temperature monitoring (NOT fire/smoke detection - no combustion) registers loss of cooling; the building could not maintain a safe operating temperature. No fire alarm - a thermal, not fire, condition.
- MITIGATION Protective thermal load-shed (de-energisation analogue): Google powers down part of the zone and limits preemptible launches to prevent machine damage. NO fire suppression discharged - no gaseous/water-mist/sprinkler activation, because there was no fire.
- IMPACT Compute Engine terminates VMs in the impacted datacenter - a small set per the public status page; ~35% of europe-west2-a VMs per The Register's report of Google's fuller incident report.
- CASCADE Recovery-procedure error widens blast radius: engineers inadvertently modify traffic routing for internal services to avoid all three europe-west2 zones rather than just the impacted europe-west2-a.
- DETECTION Google publicly acknowledges the fault: 'There has been a cooling-related failure in one of our buildings that hosts zone europe-west2-a for region europe-west2.'
- TRIGGER Oracle UK South (London): a subset of cooling infrastructure experiences an issue as a result of unseasonal temperatures - a second, independent London campus hit by the same heat event.
- DETECTION Oracle acknowledges the cooling problem at its UK South (London) data centre.
- MITIGATION Oracle protective power-down: identifying 'service infrastructure that can be safely powered down to prevent additional hardware failures' - thermal load-shed, not fire response.
- IMPACT NEGATIVE FORENSICS - no personnel evacuation reported (equipment-protection shutdown, not a life-safety event) and no emergency-services / fire-brigade call or arrival is reported; response was operator engineering teams restoring cooling and rebuilding capacity.
- RECOVERY Google cooling reported restored (~14:13 US/Pacific per The Register 2022-08-01) and capacity rebuild begins in europe-west2-a.
- RESTORED Google: 'the issue has been resolved for all affected users as of Tuesday, 2022-07-19 20:43 US/Pacific' (~14h10m from 06:33 start), though some storage issues persist requiring support contact.
- IMPACT Storage long-tail: a small number of HDD-backed Persistent Disk volumes still exhibit IO errors, with a small percentage of replicated PDs running in single-redundant mode - genuine media damage from the thermal event/abrupt power-down.
- RECOVERY Oracle data-centre cooling infrastructure reported restored and temperatures returned to normal operating levels.
- RESTORED Google official incident window closes. Status-page window: began 2022-07-19 06:33, ended 2022-07-20 21:20 US/Pacific (~38h47m to full closure); the extra tail beyond the 20:43 Jul 19 'resolved for all users' mark was storage remediation.
Root cause
Contributing factors
- RECORD AMBIENT / CLIMATE FORCING: The UK hit an all-time record 40.3C (104.5F) on 2022-07-19 - London's first day ever above 40C - pushing cooling units to and beyond their rated design envelope (The Register, 2022-07-19; design-limit framing via Tech Monitor/The Register, 2022-08-22, attributed-not-re-verified).
- DESIGN-MARGIN / WORST-CASE ENVELOPE SET TOO LOW: The plant appears to have lacked head-room for a 40C day; trade analysis contrasts operators that design and factory-test cooling 'for worst-case conditions' with the failed London plant, implying its worst-case design temperature trailed the emerging climate reality (The Register, 2022-08-22, trade analysis - not operator-confirmed).
- REDUNDANCY THAT DID NOT DELIVER (suspected common-mode / single point of failure): Google's fuller report (via The Register, 2022-08-01) described a simultaneous failure of multiple redundant cooling systems; simultaneity of 'redundant' units implies a shared driver (record ambient) or single point of failure rather than independent random failure.
- SELF-ADMITTED OPERATIONAL / RECOVERY-PROCEDURE ERROR: Engineers inadvertently modified traffic routing for internal services to avoid all three europe-west2 zones rather than just the impacted europe-west2-a, enlarging the blast radius beyond the physically affected zone (The Register, 2022-08-01; not on public status page).
- NO MAINTENANCE/INSPECTION LAPSE EVIDENCED, BUT TEST-UNDER-LOAD UNVERIFIED: No source attributes the failure to deferred maintenance or a botched service; however, the record is silent on whether the redundant units had ever been function-tested under near-design high ambient - an unverified resilience gap (disclosed uncertainty).
- SIMULTANEOUS TWO-OPERATOR EXPOSURE: Both Google (europe-west2-a) and Oracle (UK South) cooling plant failed the same day under the same heat event, indicating a regional/industry design-assumption gap rather than a single-site defect (The Register, 2022-07-19).
- ABRUPT DE-ENERGISATION AS A DAMAGE SOURCE: The protective emergency power-down itself contributed to lasting harm - replicated Persistent Disks left in single-redundant mode and HDD-backed volumes exhibiting IO errors (Google status page, 2022-07-19/20).
Correction of errors (COE)
- Raise cooling-plant worst-case design ambient and factory-test to it (climate-resilience uplift for UK/comparable sites)
- Add guardrails so zone-scoped traffic-routing mitigations cannot be applied region-wide (blast-radius containment)
- Function-test redundant cooling units under near-design high ambient; verify economiser/chiller changeover at design-max
- Restore Persistent Disk redundancy and remediate HDD-backed volumes exhibiting IO errors
- Publish equipment/plant type and maintenance-test history to enable independent post-incident learning
Lessons learnt
- CLIMATE HAS MOVED THE DESIGN TAIL: 'Once-in-a-lifetime' ambient is now design-relevant; cooling worst-case envelopes calibrated on historical UK climate under-provision for 40C+ days and must be revised upward.
- 'REDUNDANT' IS NOT 'INDEPENDENT': N+1/N+2 duplication provides no protection when all units share the same failure driver (record ambient); redundancy must be diverse and common-mode-analysed, not merely duplicated.
- BLAST RADIUS IS OFTEN OPERATIONAL, NOT PHYSICAL: The largest widening of customer impact came from a human failover-tooling error (region-wide reroute), not the cooling failure itself - recovery-tooling discipline matters as much as the plant.
- EMERGENCY DE-ENERGISATION HAS ITS OWN COST: Protective thermal shutdown left HDD-backed Persistent Disks with IO errors and replicated PDs single-redundant - the 'safe' action still damaged data durability.
- MULTI-OPERATOR SAME-DAY FAILURE IS AN INDUSTRY SIGNAL: Two independent London operators failing on the same heat event points to a shared design-assumption gap across the sector, not an isolated site defect.
- DISCLOSURE WAS PARTIAL: Neither operator released equipment make/model/age, refrigerant, plant type, or maintenance/test history - limiting external learning and independent verification.
Improvements & remediation
- DESIGN HEAD-ROOM: Raise the cooling-plant worst-case design ambient and factory-test plant to it, so cooling retains capacity on 40C+ days rather than being 'required to operate above design limits'; UK worst-case envelopes set on historical climate are now obsolete.
- REDUNDANCY DIVERSITY (Safety-critical): Audit London (and comparable) plant for common-mode failure paths and single points of failure so that nominally 'redundant' units cannot be defeated simultaneously by a shared driver such as record ambient; add independent/diverse cooling paths, not just duplicated identical units.
- MAINTENANCE / FUNCTION-TESTING: Institute periodic function-testing of standby and redundant cooling units under near-design HIGH-ambient load (not only cool-day proof runs); verify chiller/economiser changeover and setpoints hold at design-max wet-bulb, and log test-under-load results.
- THERMAL-PROTECTION SAFETY RUNBOOK: Formalise a graceful, staged thermal load-shed runbook (migrate/drain workloads before hard VM termination where time permits) and validate BMS high-temperature alarm thresholds and shutdown setpoints to buy migration time and reduce abrupt-power-down data damage.
- RECOVERY-TOOLING GUARDRAILS: Add blast-radius containment to traffic-routing/failover tooling so a zone-scoped mitigation cannot be inadvertently applied region-wide (the europe-west2 all-three-zone reroute error).
- STORAGE DURABILITY: Harden Persistent Disk against abrupt de-energisation so emergency shutdown does not leave replicated volumes single-redundant or HDD-backed volumes with IO errors; prioritise redundancy-restore automation post-event.
- CUSTOMER RESILIENCE GUIDANCE: Reinforce multi-zone/multi-region architecture guidance, treating a single zone/building as a failure domain that can be lost to a facility-level thermal event.
Comprehensive analysis
What actually happened - and what did not
On 2022-07-19, the UK's hottest day on record (40.3C), cooling failed at two independent London cloud campuses - Google's europe-west2-a and Oracle's UK South. This was a thermal cooling-capacity collapse, not a fire. There was no ignition, no combustion, no smoke/heat detection, no suppression discharge, no evacuation, and no emergency-services response reported by any source. The correct analogues are environmental/BMS temperature monitoring in place of fire detection, and deliberate thermal load-shedding / de-energisation in place of fire suppression. Google confirmed 'a cooling related failure in one of our buildings that hosts zone europe-west2-a', and 'powered down part of the zone' to prevent machine damage.
The latent root: design margin versus a changing climate
The units were, per trade reporting, 'required to operate above their design limits' - a design-margin gap rather than a maintenance failure. The plant's worst-case ambient envelope appears to have been set below the emerging UK climate reality; London had never before exceeded 40C. This reframes the event from an equipment fault to a climate-resilience shortfall: cooling designed and factory-tested for historical worst-case conditions is now under-provisioned. (The 'above design limits' framing is trade analysis via Tech Monitor/The Register and was not operator-confirmed or re-verified this pass.)
Redundancy that did not deliver
Google's fuller report (via The Register, 2022-08-01) described a simultaneous failure of multiple redundant cooling systems. Simultaneity among 'redundant' units is the tell of a common-mode driver - here the shared record ambient - or a single point of failure. Genuine N+1/N+2 redundancy should not permit backup-capable units to fail together; that they did indicates the redundancy was duplicated rather than diverse and independent. This is the single most important engineering lesson of the event.
Self-inflicted blast radius: the failover-tooling error
The physical impact was confined to one building/hall of europe-west2-a (europe-west2-b and -c were explicitly unaffected for VMs). The damaging secondary spread was operational: engineers inadvertently rerouted internal-service traffic away from all three europe-west2 zones rather than just the impacted one, widening customer impact. This procedural error - not the cooling failure - drove much of the customer-visible blast radius, underscoring recovery-tooling discipline. (Reported by The Register; not on the public status page.)
Storage long-tail and the cost of emergency shutdown
Protective de-energisation is not free. Google's status page confirms 'a small percentage of replicated Persistent Disk devices are now running in single redundant mode' and 'a small number of HDD backed Persistent Disk volumes are still experiencing impact and will exhibit IO errors.' The abrupt power-down inflicted genuine storage-media damage and left durability degraded, extending the official incident window (06:33 Jul 19 to 21:20 Jul 20 US/Pacific) well past the 'resolved for all users' mark at 20:43 Jul 19.
Two operators, one heat event, and the disclosure gap
Both Google and Oracle cooling plant failed the same day under the same heat - a sector-level design-assumption signal, not a one-off. Yet disclosure was thin: neither operator released equipment make/model/age, refrigerant chemistry, plant type (free-cooling economiser vs mechanical DX/chilled-water), maintenance/test history, MW lost, rack counts, or cost. No maintenance or inspection lapse is evidenced anywhere; equally, whether the redundant units had ever been function-tested under near-design ambient is unknown. These remain disclosed-uncertainty points, not findings.
Technical deep-dive
References & provenance
- vendor-status Google Cloud Service Health - Incident XVq5om2XEDSqLtJZUvcH (Compute Engine cooling failure, europe-west2-a)“A small percentage of replicated Persistent Disk devices are now running in single redundant mode.”https://status.cloud.google.com/incidents/XVq5om2XEDSqLtJZUvcH
- press The Register - 'Heat wave knocks out Google, Oracle clouds in London' (2022-07-19)“As a result of unseasonal temperatures in the region, a subset of cooling infrastructure within the UK South (London) Data Centre has experienced an issue.”https://www.theregister.com/2022/07/19/google_oracle_cloud/
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).