British Airways Boadicea House Power Surge: A 15-Minute Fault That Grounded Heathrow and Gatwick for Days
On the morning of Saturday 27 May 2017, an engineer working inside British Airways' Boadicea House data centre in West London disconnected a power supply and then reconnected it in an uncontrolled, unplanned manner. The resulting surge physically damaged servers and bypassed the site's backup generators and batteries. Although the BoHo site itself was without power for only about 15 minutes, the physical damage plus a failed failover to the secondary data centre escalated a brief electrical fault into a multi-day service outage. British Airways grounded all flights at Heathrow and Gatwick, affecting roughly 75,000 passengers and as many as 672 flights. IAG later confirmed a loss of about £58m. No formal operator postmortem was ever published; BA sued its facilities supplier CBRE and the dispute settled in February 2019 with no admission of liability.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- British Airways (IAG); data-centre facility management by CBRE
- Data center
- Boadicea House (BoHo) data centre, West London
- Location
- London, United Kingdom
- Date
- 2017-05-27
Impact & scale
- Users affected
- ~75,000 passengers affected; as many as 672 flights grounded across Heathrow and Gatwick
- Financial
- ~£58m (~$74.6m), IAG-confirmed July 2017 (compensation + lost business); higher press figures ($100m, up to £150m) unreconciled
- Scope
- Tier-1 / critical business outage (dual-hub airline grounding)
- Passenger check-in
- Baggage handling
- Operational communications / core airline systems
- Online booking-data integrity (stale flight data)
Impact data & metrics
| UPS backup rating | Up to 2.4 MW (Socomec Smart Powerport) |
| IT footprint | ~500 cabinets across six halls |
| Data centres | 2, running active:active |
| UPS age at incident | ~3 years (replaced c.2014) |
| Incident start | 09:30, Saturday 27 May 2017 |
| Distance to Heathrow runways | No more than ~1 mile |
| Hypothesised surge voltage | ~480 V vs ~240 V nominal (conjecture, unconfirmed) |
| Passengers affected | ~75,000 (press-reported) |
| Estimated cost to IAG | ~£80m (press-reported) |
Magnitude profile
Blast radius is elevated because the outage grounded two of the UK's busiest airport hubs simultaneously and defeated a designed active/active redundancy pair, even though impact was contained to a single airline. Duration and financial scores reflect a ~2-day grounding and an IAG-confirmed ~£58m loss; the striking feature is the gap between a ~15-minute physical power fault and a multi-day recovery.
Sequence of events (SOE)
- TRIGGER BA replaces Boadicea House UPSes with Socomec 'Smart Powerport' equipment specifiable up to 2.4 MW of backup, backed by battery strings and standby generators; two data centres run active:active, ~500 cabinets across six halls, within ~1 mile of Heathrow's eastern-end runways.
- TRIGGER The UPS feeding the core Boadicea House data centre is manually 'over-ridden' during on-site work at the CBRE-managed facility.
- CASCADE The override causes 'the total immediate loss of power to the facility, bypassing the backup generators and batteries' — the redundant supply path is stranded rather than picking up the load.
- DETECTION NOT APPLICABLE / no fire: no smoke, heat, VESDA or aspirating detection activated because no combustion occurred; the event was electrical over-voltage, detected operationally as total loss of the data hall, not by fire sensors.
- TRIGGER Power is 'turned back on in an unplanned and uncontrolled fashion' — an unsequenced re-energisation with no staged load ramp or downstream verification.
- IMPACT The uncontrolled re-energisation produces a damaging over-voltage surge (hypothesised battery+generator briefly in series, ~480 V vs ~240 V nominal) that physically damages servers and infrastructure.
- MITIGATION No gaseous or water-mist SUPPRESSION activated — NOT APPLICABLE: there was no fire to suppress; the damage vector was electrical, so suppression systems had no role and none is reported.
- CASCADE Active:active redundancy fails to contain the event; the peer data centre also collapses — reported (unconfirmed) mechanism is that corrupted synchronised data left the failover DC 'populated with enough bad data to crash all the systems depending on it.'
- IMPACT BA's global operations go down; flights cancelled and passengers stranded at Heathrow Terminal 5 and Gatwick on a bank-holiday Saturday. No occupant evacuation of Boadicea House occurred (operational, not life-safety, impact).
- MITIGATION No London Fire Brigade or ambulance dispatch to Boadicea House — NOT APPLICABLE: no fire or medical emergency at the facility; the emergency-services layer has no entries in the public record.
- MITIGATION De-energisation/isolation and recovery: because hardware was physically damaged (not merely powered down), recovery requires rebuild/restore of damaged infrastructure rather than a clean seconds-long failover.
- DETECTION BA chief executive Alex Cruz publicly attributes the outage to a major 'power surge' at 09:30 on Saturday 27 May that caused systems to 'collapse.'
- DETECTION National Grid states there was no transmission-network problem in the Heathrow area that weekend; SSE states any surge 'could have taken place at the customer side of the meter' — locating the fault inside BA's facility.
- RECOVERY BA progressively restores IT systems and flight operations over the following days as damaged infrastructure is rebuilt/restored; service normalises after the bank-holiday weekend disruption.
- RESTORED Leaked internal IAG email from Bill Francis (Head of Group IT) detailing the UPS override and uncontrolled reconnection is reported; CBRE confirms it manages the site, 'fully support[s]' the investigation and says 'no determination has been made yet regarding the cause'; the Daily Mail names a CBRE Global Workplace Solutions contractor.
- RESTORED No adjudicated root cause is ever published: the BA/IAG-commissioned independent investigation report is not publicly released and the BA/CBRE civil claim is reportedly settled privately (not verified this pass).
Root cause
Contributing factors
- MAINTENANCE / MOP FAILURE (proximate): the UPS was manually 'over-ridden' and then power was restored 'in an unplanned and uncontrolled fashion' — the opposite of a controlled, sequenced, load-verified re-energisation. No authorised method-of-procedure or work permit for the Saturday work has ever been made public, so whether a valid MOP existed or was breached is unestablished (leaked IAG email via The Register, 2017-06-02).
- REDUNDANCY BYPASSED RATHER THAN TRANSFERRED: the override caused 'the total immediate loss of power to the facility, bypassing the backup generators and batteries.' The N+ backup path (generators + battery strings) that should have carried the critical load was stranded instead of picking it up (leaked IAG email via The Register).
- TIGHTLY-COUPLED ACTIVE:ACTIVE DESIGN: two data centres synchronised live state, so a fault at Boadicea House was not isolated — the reported (unconfirmed) mechanism is that corrupted synchronised data left the failover site 'populated with enough bad data to crash all the systems depending on it,' turning redundancy into a shared failure path (The Register, source-attributed theory).
- UNCONTROLLED RE-ENERGISATION WITH NO STAGED RAMP: reconnecting power without sequencing sources or verifying downstream state produced a damaging over-voltage (hypothesised battery+generator briefly in series, ~480 V vs ~240 V nominal) that physically destroyed hardware, forcing rebuild/restore rather than clean failover (The Register, expert hypothesis — unconfirmed).
- SPLIT OPERATIONAL ACCOUNTABILITY BA/CBRE: the facility was managed by CBRE Global Workplace Solutions under contract to BA; the contractor allegedly involved was named only by the Daily Mail (via The Register) while CBRE said 'no determination has been made yet regarding the cause,' indicating unclear ownership of the switching operation and its controls.
- PHYSICAL HARDWARE DAMAGE (not just outage): the surge 'created physical damage to the system,' meaning recovery required replacing/rebuilding damaged infrastructure rather than a seconds-long failover — a major driver of the multi-day operational impact (leaked IAG email via The Register).
- TIMING / CHANGE-CONTROL EXPOSURE: live critical-power switching was performed on a bank-holiday Saturday at the start of a peak travel weekend, maximising passenger impact when a change-freeze on peak windows would have reduced exposure (Alex Cruz confirmed 09:30 Sat 27 May start, via The Register).
Correction of errors (COE)
- Publish and enforce an approved MOP/SOP/EOP for all UPS bypass and re-energisation operations on critical power
- Independent investigation to establish adjudicated root cause
- Re-architect active:active replication to gate on data integrity and prevent corrupt-state propagation to the failover site
- Clarify BA/CBRE RACI for critical-power switching and its interlock/protection controls
- Civil claim BA/IAG vs CBRE over the outage
Lessons learnt
- Redundancy is only as good as the switching procedure around it: a single uncontrolled human switching action defeated the generators, batteries and a twin active:active data centre all at once.
- Tight active:active coupling can convert a local fault into a shared-mode failure — synchronisation that spreads good state will just as faithfully spread corrupted state to the site meant to be your safety net.
- The lasting damage came from RE-ENERGISATION, not the outage: the most dangerous moment in critical-power work is turning power back on, which must be staged, phase-synchronised and downstream-verified — never 'unplanned and uncontrolled.'
- Split operational accountability between an asset owner (BA) and a facilities-management contractor (CBRE) blurs who owns critical switching and its interlocks; that ambiguity is itself a latent hazard.
- Relying on external forensics to learn from your own incident is fragile: with the investigation report unreleased and the claim reportedly settled privately, the public lessons here rest almost entirely on one leaked email and press reporting.
Improvements & remediation
- MAINTENANCE / MOP: Mandate a written, peer-reviewed Method-of-Procedure (with matching SOP/EOP) for every UPS bypass and re-energisation on critical power. No live switching without an approved, staged, load-verified sequence and a second competent person independently verifying each step and downstream breaker/PDU state before the next.
- MAINTENANCE / ELECTRICAL: Enforce true make-before-break maintenance bypass so the critical load transfers to raw mains or generator BEFORE any UPS module work — never a break-before-make operation that drops the load and simultaneously strands generators and batteries as happened here.
- SAFETY / PROTECTION: Fit and routinely test source-transfer interlocks and sync-check relays that physically prevent paralleling battery and generator supplies out-of-phase or in series; add rack/PDU-level over-voltage protection and monitoring so an operator switching error cannot deliver double voltage to server PSUs.
- RESILIENCE / ARCHITECTURE: Loosen active:active coupling with data-integrity gating on replication so an uncommanded shutdown at one site cannot propagate corrupted state to its twin; validate failover with regular full-scale DR tests that deliberately include dirty-shutdown and hardware-damage scenarios, not just clean cutovers.
- GOVERNANCE / ACCOUNTABILITY: Publish a clear BA/CBRE RACI for critical-power switching; require BA change-authority sign-off, and impose change-freeze windows over peak/bank-holiday travel periods before any live critical-power work.
- RECOVERY / RUNBOOKS: Maintain out-of-band, geographically-independent recovery capability and rehearsed rebuild/restore runbooks that assume physical hardware destruction can make failover impossible and force a multi-day rebuild.
Comprehensive analysis
Not a fire — an electrical over-voltage event
The single most important forensic correction: this was NOT a fire and must never be framed as a 'battery fire.' There was no ignition, smoke, heat, VESDA/aspirating detection, gaseous or water-mist suppression, evacuation, or emergency-services response, because none occurred and none was applicable. The batteries were bypassed, not thermally involved. The damaging agent was an uncontrolled electrical over-voltage transient generated inside BA's own switchgear. The classic fire-forensic layers are genuinely absent from the record — that is the correct physics, not a data gap to be filled.
Proximate mechanism: override then uncontrolled re-energisation
Per the leaked IAG email (Bill Francis, via The Register), a UPS to a core Heathrow data centre 'was over-ridden,' causing 'the total immediate loss of power to the facility, bypassing the backup generators and batteries.' Minutes later it 'was turned back on in an unplanned and uncontrolled fashion, which created physical damage to the system.' The damage came from the restoration, not the outage: a hypothesised momentary series connection of battery and generator delivering ~480 V versus ~240 V nominal — a figure that is press-reported conjecture, not measured fact.
How redundancy propagated instead of contained
Two data centres ran active:active, ~500 cabinets across six halls, both within ~1 mile of Heathrow. Continuous state synchronisation means either site can take the whole load — and that corrupt state can replicate. The Register's source suggested the uncommanded shutdown synchronised corrupted data across, leaving the failover DC 'populated with enough bad data to crash all the systems.' Because BoHo hardware was physically destroyed, recovery was a multi-day rebuild/restore, not a seconds-long cutover — turning a few-minutes electrical event into a bank-holiday operational collapse.
Origin located inside BA's fence
Two independent utility data points bound the origin. National Grid confirmed no transmission-network problem in the Heathrow area that weekend. SSE stated any 'power surge that BA is referring to could have taken place at the customer side of the meter.' Together they place the transient inside BA's own infrastructure — consistent with an operator-induced switching event under CBRE facilities management rather than a utility disturbance.
What remains unproven and non-public
Every deep claim traces to one primary chain: the leaked IAG email as reported by The Register, plus attributed CBRE/SSE/National Grid statements. No adjudicated root cause was ever published — the independent report was not released and the BA/CBRE claim was reportedly settled privately. Unknown/disputed: who physically switched (a CBRE contractor named only by the Daily Mail); the true surge magnitude (480 V series theory is conjecture); why the peer site failed (data-corruption theory unverified); battery chemistry and exact equipment age; and whether a valid MOP existed.
Technical deep-dive
References & provenance
- press British Airways IT failure caused by 'uncontrolled return of power' (The Guardian, reproducing BA's official statement verbatim)“There was a loss of power to the UK data centre which was compounded by the uncontrolled return of power which caused a power surge taking out our IT systems.”https://www.theguardian.com/business/2017/may/31/ba-it-shutdown-caused-by-uncontrolled-return-of-power-after-outage
- press Inquiry to determine whether BA outage was human error (DatacenterDynamics)“The IT equipment was powered up in an 'uncontrolled fashion,' causing a surge and 'catastrophic physical damage' to the airline's servers, a source told The Telegraph.”https://www.datacenterdynamics.com/en/news/inquiry-to-determine-whether-ba-outage-was-human-error/
- press British Airways outage, like most data center outages, was caused by humans“There was a loss of power to the U.K. data center, which was compounded by the uncontrolled return of power, which caused a power surge taking out our IT systems.”https://www.networkworld.com/article/963711/british-airways-outage-like-most-data-center-outages-was-caused-by-humans.html
- press British Airways sues CBRE over IT meltdown“unauthorised reconnection of servers after they had been accidentally unplugged from the power supply by an engineer”https://www.ch-aviation.com/news/73249-british-airways-sues-cbre-over-it-meltdown
- news BA IT chaos caused by human error, boss says (leaked IAG email + Walsh)“it was turned back on in an unplanned and uncontrolled fashion, which created physical damage to the systems”https://www.bbc.com/news/business-40159202
- news British Airways: What went wrong? (impact and EU261 schedule)“One expert predicted disruption lasting 'three or four days.'”https://www.bbc.com/news/uk-40069865
- analysis British Airways' data centre configuration (engineering hypotheses)“That would result in the data centre's servers being fed 480v instead of 240v, causing a literal meltdown.”https://www.theregister.com/2017/06/02/british_airways_data_centre_configuration/
- news BA and CBRE settle dispute over 2017 data center outage“British Airways and CBRE are pleased to have reached agreement (with no admission as to liability)”https://www.datacenterdynamics.com/en/news/ba-and-cbre-settle-dispute-over-2017-data-center-outage/
- news BA to sue CBRE over May Bank Holiday datacentre outage“BA is understood to be taking legal action”https://www.computerweekly.com/news/252452976/BA-to-sue-CBRE-over-May-Bank-Holiday-datacentre-outage
- news British Airways $100M Outage Caused by Worker Pulling Wrong Plug“human error was to blame for the meltdown that stranded tens of thousands”https://www.nbcnews.com/business/travel/british-airways-100m-outage-was-caused-worker-pulling-wrong-plug-n767476
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.