← All incidents
Incident dossier · Rank #31

British Airways Boadicea House Power Surge: A 15-Minute Fault That Grounded Heathrow and Gatwick for Days

British Airways (IAG); data-centre facility management by CBRE 2017-05-27 48h 0m core impact PowerHuman error

On the morning of Saturday 27 May 2017, an engineer working inside British Airways' Boadicea House data centre in West London disconnected a power supply and then reconnected it in an uncontrolled, unplanned manner. The resulting surge physically damaged servers and bypassed the site's backup generators and batteries. Although the BoHo site itself was without power for only about 15 minutes, the physical damage plus a failed failover to the secondary data centre escalated a brief electrical fault into a multi-day service outage. British Airways grounded all flights at Heathrow and Gatwick, affecting roughly 75,000 passengers and as many as 672 flights. IAG later confirmed a loss of about £58m. No formal operator postmortem was ever published; BA sued its facilities supplier CBRE and the dispute settled in February 2019 with no admission of liability.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Power (2017-05-27)Trigger · Power2017-05-272017-05-27Primary fault at British Airways (IAG); data-centre facility management by CBRE — Boadicea House (BoHo) data centre, West LondonBritish AirwaysBoadicea House (BoHo) data centre, West LondonBoadicea HouseDownstream service degraded by the fault: Passenger check-inPassenger check-inDownstream service degraded by the fault: Baggage handlingBaggage handlingDownstream service degraded by the fault: Operational communications / core airline systemsOperational communications /core…Downstream service degraded by the fault: Online booking-data integrity (stale flight data)Online booking-data integrity

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
British Airways (IAG); data-centre facility management by CBRE
Data center
Boadicea House (BoHo) data centre, West London
Location
London, United Kingdom
Date
2017-05-27

Impact & scale

Users affected
~75,000 passengers affected; as many as 672 flights grounded across Heathrow and Gatwick
Financial
~£58m (~$74.6m), IAG-confirmed July 2017 (compensation + lost business); higher press figures ($100m, up to £150m) unreconciled
Scope
Tier-1 / critical business outage (dual-hub airline grounding)
Services / systems down
  • Passenger check-in
  • Baggage handling
  • Operational communications / core airline systems
  • Online booking-data integrity (stale flight data)

Impact data & metrics

UPS backup ratingUp to 2.4 MW (Socomec Smart Powerport)
IT footprint~500 cabinets across six halls
Data centres2, running active:active
UPS age at incident~3 years (replaced c.2014)
Incident start09:30, Saturday 27 May 2017
Distance to Heathrow runwaysNo more than ~1 mile
Hypothesised surge voltage~480 V vs ~240 V nominal (conjecture, unconfirmed)
Passengers affected~75,000 (press-reported)
Estimated cost to IAG~£80m (press-reported)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 6Users affected (0–10) — breadth of the user/customer population impacted. — scored 6/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 6Outage duration (0–10) — how long service was degraded/down. — scored 6/10.Blast 7Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 7/10.
Magnitude 6.3 = blast 7×0.35 + users 6×0.25 + financial 6×0.20 + duration 6×0.20 (sub-scores 0–10 · weighted composite)

Blast radius is elevated because the outage grounded two of the UK's busiest airport hubs simultaneously and defeated a designed active/active redundancy pair, even though impact was contained to a single airline. Duration and financial scores reflect a ~2-day grounding and an IAG-confirmed ~£58m loss; the striking feature is the gap between a ~15-minute physical power fault and a multi-day recovery.

Sequence of events (SOE)

Phased sequence of eventsc. 2014 (~3 years prior) · TRIGGER — BA replaces Boadicea House UPSes with Socomec 'Smart Powerport' equipment specifiable up to 2.4 MW of backup, backed by battery strings and standby generators; two data centres run active:active, ~500 cabinets across six halls, within ~1 mile of Heathrow's eastern-end runways.TRIGGERc. 2014 (~3 year2017-05-27, ~09:30 (Saturday) · TRIGGER — The UPS feeding the core Boadicea House data centre is manually 'over-ridden' during on-site work at the CBRE-managed facility.TRIGGER2017-05-27, ~09:2017-05-27, ~09:30 · CASCADE — The override causes 'the total immediate loss of power to the facility, bypassing the backup generators and batteries' — the redundant supply path is stranded rather than picking up the load.CASCADE2017-05-27, ~09:2017-05-27, ~09:30 · DETECTION — NOT APPLICABLE / no fire: no smoke, heat, VESDA or aspirating detection activated because no combustion occurred; the event was electrical over-voltage, detected operationally as total loss of the data hall, not by fire sensors.DETECTION2017-05-27, ~09:2017-05-27, ~a few minutes after 09:30 · TRIGGER — Power is 'turned back on in an unplanned and uncontrolled fashion' — an unsequenced re-energisation with no staged load ramp or downstream verification.TRIGGER2017-05-27, ~a f2017-05-27, ~a few minutes after 09:30 · IMPACT — The uncontrolled re-energisation produces a damaging over-voltage surge (hypothesised battery+generator briefly in series, ~480 V vs ~240 V nominal) that physically damages servers and infrastructure.IMPACT2017-05-27, ~a f2017-05-27, morning · MITIGATION — No gaseous or water-mist SUPPRESSION activated — NOT APPLICABLE: there was no fire to suppress; the damage vector was electrical, so suppression systems had no role and none is reported.MITIGATION2017-05-27, morn2017-05-27, morning · CASCADE — Active:active redundancy fails to contain the event; the peer data centre also collapses — reported (unconfirmed) mechanism is that corrupted synchronised data left the failover DC 'populated with enough bad data to crash all the systems depending on it.'CASCADE2017-05-27, morn2017-05-27, morning · IMPACT — BA's global operations go down; flights cancelled and passengers stranded at Heathrow Terminal 5 and Gatwick on a bank-holiday Saturday. No occupant evacuation of Boadicea House occurred (operational, not life-safety, impact).IMPACT2017-05-27, morn2017-05-27, morning · MITIGATION — No London Fire Brigade or ambulance dispatch to Boadicea House — NOT APPLICABLE: no fire or medical emergency at the facility; the emergency-services layer has no entries in the public record.MITIGATION2017-05-27, morn2017-05-27, day · MITIGATION — De-energisation/isolation and recovery: because hardware was physically damaged (not merely powered down), recovery requires rebuild/restore of damaged infrastructure rather than a clean seconds-long failover.MITIGATION2017-05-27, day2017-05-27 / 28 · DETECTION — BA chief executive Alex Cruz publicly attributes the outage to a major 'power surge' at 09:30 on Saturday 27 May that caused systems to 'collapse.'DETECTION2017-05-27 / 282017-05-27 / 28 · DETECTION — National Grid states there was no transmission-network problem in the Heathrow area that weekend; SSE states any surge 'could have taken place at the customer side of the meter' — locating the fault inside BA's facility.DETECTION2017-05-27 / 282017-05-28 onward · RECOVERY — BA progressively restores IT systems and flight operations over the following days as damaged infrastructure is rebuilt/restored; service normalises after the bank-holiday weekend disruption.RECOVERY2017-05-28 onwar2017-06-02 · RESTORED — Leaked internal IAG email from Bill Francis (Head of Group IT) detailing the UPS override and uncontrolled reconnection is reported; CBRE confirms it manages the site, 'fully support[s]' the investigation and says 'no determination has been made yet regarding the cause'; the Daily Mail names a CBRE Global Workplace Solutions contractor.RESTORED2017-06-02After 2017 · RESTORED — No adjudicated root cause is ever published: the BA/IAG-commissioned independent investigation report is not publicly released and the BA/CBRE civil claim is reportedly settled privately (not verified this pass).RESTOREDAfter 2017

Root cause

SPECIFIC MECHANISM (not fire — electrical over-voltage). There was no ignition, no combustion, no flame. The damaging agent was an uncontrolled electrical over-voltage transient — a power surge generated inside British Airways' own facility (Boadicea House / "BoHo", a core data centre within ~1 mile of Heathrow's eastern-end runways). The central equipment was the site Uninterruptible Power Supply: reported by The Register (2017-06-02) as Socomec "Smart Powerport" units, replaced ~3 years before the 27 May 2017 incident (i.e. c.2014), rated to supply up to 2.4 MW of backup, backed by battery strings and standby generators. Battery chemistry was NEVER stated in the record; era-typical UPS strings are VRLA lead-acid, but that is inference, not sourced fact. Crucially the batteries were BYPASSED, not thermally involved — so there is NO battery thermal-runaway or fire mechanism here, and any framing of this as a "battery fire" would be false. PROXIMATE FAILURE MECHANISM (sourced). Per the leaked internal IAG email from Bill Francis (Head of Group IT), an Uninterruptible Power Supply to a core data centre at Heathrow "was over-ridden on Saturday morning." The email states this "resulted in the total immediate loss of power to the facility, bypassing the backup generators and batteries." Then "after a few minutes of this shutdown of power, it was turned back on in an unplanned and uncontrolled fashion, which created physical damage to the system." Sequence: (1) manual override/bypass of the UPS that dropped the entire critical load and defeated — rather than transferred onto — the generator+battery redundancy; then (2) an uncontrolled, unsequenced re-energisation that produced a damaging over-voltage which physically destroyed IT hardware. The leading hypothesis, reported by The Register, is that improper switching "could have briefly connected both [battery and generator] in series, delivering 480 volts instead of the normal 240 volts to server racks." This ~480 V vs ~240 V figure is conjecture reported in the press, NOT a confirmed measured engineering finding; BA never published fault data. External corroboration that the fault originated inside BA's fence: National Grid confirmed no transmission-network problem in the Heathrow area that weekend, and Scottish and Southern (SSE) stated "the power surge that BA is referring to could have taken place at the customer side of the meter." LATENT ROOT — design + redundancy. Design intent was full resilience: two data centres running active:active, ~500 cabinets across six halls, with UPS+battery+generator backup rated to 2.4 MW. That redundancy did not contain the event; it propagated it. The reported (unconfirmed, attributed to The Register's "source") mechanism is that an uncommanded shutdown caused corrupted data to be synchronised between the two sites as BoHo died, so the failover DC "could have been populated with enough bad data to crash all the systems depending on it" — i.e. tight active:active coupling let a fault at one site contaminate its twin, converting redundancy from a safety margin into a shared failure path. The physical over-voltage also meant recovery was not a clean failover but a hardware rebuild/restore, which is why the outage ran for days rather than seconds. LATENT ROOT — maintenance / method-of-procedure. The implicated (never adjudicated) lapse is a work-procedure / method-of-procedure (MOP) failure. The proximate human act — over-riding the UPS and then restoring power "in an unplanned and uncontrolled fashion" — is the antithesis of a controlled, staged, load-verified re-energisation. It implies either no valid MOP for the Saturday work, or a MOP that was breached, plus an operation that bypassed generators AND batteries instead of transferring load onto them. The facility was managed by CBRE Global Workplace Solutions under contract to BA; the Daily Mail (reported via The Register) "fingered a contractor from CBRE Global Workplace Solutions as the culprit," but CBRE publicly said only that it is "the manager of the facility" and "fully support[s]" BA's investigation and that "no determination has been made yet regarding the cause." No maintenance log, testing record or work permit for that Saturday has ever been made public. BA/IAG commissioned an independent investigation that was never publicly released, and the BA/CBRE civil claim was reportedly settled privately (not verified this pass), so the official root cause remains legally unresolved and largely non-public.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Not a fire — an electrical over-voltage event

The single most important forensic correction: this was NOT a fire and must never be framed as a 'battery fire.' There was no ignition, smoke, heat, VESDA/aspirating detection, gaseous or water-mist suppression, evacuation, or emergency-services response, because none occurred and none was applicable. The batteries were bypassed, not thermally involved. The damaging agent was an uncontrolled electrical over-voltage transient generated inside BA's own switchgear. The classic fire-forensic layers are genuinely absent from the record — that is the correct physics, not a data gap to be filled.

Proximate mechanism: override then uncontrolled re-energisation

Per the leaked IAG email (Bill Francis, via The Register), a UPS to a core Heathrow data centre 'was over-ridden,' causing 'the total immediate loss of power to the facility, bypassing the backup generators and batteries.' Minutes later it 'was turned back on in an unplanned and uncontrolled fashion, which created physical damage to the system.' The damage came from the restoration, not the outage: a hypothesised momentary series connection of battery and generator delivering ~480 V versus ~240 V nominal — a figure that is press-reported conjecture, not measured fact.

How redundancy propagated instead of contained

Two data centres ran active:active, ~500 cabinets across six halls, both within ~1 mile of Heathrow. Continuous state synchronisation means either site can take the whole load — and that corrupt state can replicate. The Register's source suggested the uncommanded shutdown synchronised corrupted data across, leaving the failover DC 'populated with enough bad data to crash all the systems.' Because BoHo hardware was physically destroyed, recovery was a multi-day rebuild/restore, not a seconds-long cutover — turning a few-minutes electrical event into a bank-holiday operational collapse.

Origin located inside BA's fence

Two independent utility data points bound the origin. National Grid confirmed no transmission-network problem in the Heathrow area that weekend. SSE stated any 'power surge that BA is referring to could have taken place at the customer side of the meter.' Together they place the transient inside BA's own infrastructure — consistent with an operator-induced switching event under CBRE facilities management rather than a utility disturbance.

What remains unproven and non-public

Every deep claim traces to one primary chain: the leaked IAG email as reported by The Register, plus attributed CBRE/SSE/National Grid statements. No adjudicated root cause was ever published — the independent report was not released and the BA/CBRE claim was reportedly settled privately. Unknown/disputed: who physically switched (a CBRE contractor named only by the Daily Mail); the true surge magnitude (480 V series theory is conjecture); why the peer site failed (data-corruption theory unverified); battery chemistry and exact equipment age; and whether a valid MOP existed.

Technical deep-dive

This incident is a textbook case of redundancy defeating itself, and it is important to state up front — with critical honesty — what it was NOT. It was not a fire. No smoke, heat, VESDA/aspirating detection, gaseous suppression (inert gas / FM-200 / Novec), water mist, evacuation or emergency-services response is reported, because none occurred and none was applicable. The classic fire-forensic layers simply do not exist in the public record for this event; treating them as "missing data" would misrepresent the physics. The damage vector was electrical over-voltage, not combustion. The electrical architecture (The Register, 2017-06-02, quoting the leaked IAG email from Bill Francis). Boadicea House was fed through a site UPS reported as Socomec "Smart Powerport," replaced ~3 years prior and specifiable up to 2.4 MW of backup, with battery strings for ride-through and standby generators for extended outages. This is a conventional double-conversion topology: mains normally feeds the load through the UPS while charging batteries; on mains loss the batteries carry the load for the seconds/minutes it takes generators to start and assume it. Correctly operated, a maintenance bypass transfers the critical load to raw mains (or generator) so the UPS module can be worked on WITHOUT dropping the racks. What the record describes is the opposite: the UPS was "over-ridden," and the result was "the total immediate loss of power to the facility, bypassing the backup generators and batteries." A bypass that drops the load and simultaneously strands both backup sources is not a make-before-break maintenance bypass functioning as designed — it is a break-before-anything event. The surge mechanism. The physical damage came not from the power loss but from the restoration: "after a few minutes of this shutdown of power, it was turned back on in an unplanned and uncontrolled fashion, which created physical damage to the system." A controlled re-energisation ramps load in stages, verifies downstream breaker/PDU state, and synchronises sources before paralleling them. The reported hypothesis is that improper switching momentarily paralleled battery and generator supplies in series, roughly doubling rack voltage to ~480 V versus a nominal ~240 V — enough to destroy server power-supply units outright. This series-connection / 480 V figure must be flagged as conjecture reported by The Register, not a measured, published finding; BA never released fault waveforms or metering. Two external data points bound the origin: National Grid reported no transmission-network fault near Heathrow, and SSE stated any surge "could have taken place at the customer side of the meter." Together they place the transient inside BA's own switchgear — consistent with an operator-induced event rather than a utility disturbance. Why redundancy propagated the fault. The two data centres ran active:active, ~500 cabinets across six halls, both within ~1 mile of Heathrow's eastern-end runways. Active:active means both sites carry live production and continuously synchronise state. The intended benefit — either site can take the whole load — is also the intended hazard: a corrupt state can replicate. The Register's source suggested the uncommanded shutdown "may have caused corrupted data to be synchronised between the two as BoHo died," populating the failover DC "with enough bad data to crash all the systems depending on it." Combined with the fact that Boadicea House hardware was physically damaged (not merely powered down), recovery could not be a clean cutover; it required rebuilding/restoring damaged infrastructure, which is why a few-minutes electrical event became a multi-day operational collapse cancelling flights at Heathrow Terminal 5 and Gatwick on a bank-holiday Saturday. What remains unproven. Every deep claim above traces to a single primary chain — the leaked IAG email as reported by The Register, plus attributed CBRE/SSE/National Grid statements. No adjudicated root cause was ever published: the independent investigation report was not released and the BA/CBRE claim was reportedly settled privately. Disputed or unknown: who physically performed the override/reconnection (a CBRE contractor named only by the Daily Mail, never officially confirmed); the true surge magnitude and mechanism (the 480 V series theory is conjecture); why the peer site failed (data-corruption theory attributed to an unnamed source, unverified); exact battery chemistry, UPS variant, and true equipment age beyond "replaced three years ago"; whether a valid MOP existed and how it was breached; and the precise minute-by-minute internal timeline beyond the 09:30 start and "a few minutes later" reconnection.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home