← All incidents
Incident dossier · Rank #18

Equinix LD8 Docklands UPS Failure: ~17.5-Hour Power Outage Disrupts LINX and ~150 Members

Equinix 2020-08-18 17h 30m core impact Power

On 18 August 2020 a faulty UPS system at Equinix's LD8 Docklands data centre — home to the London Internet Exchange (LINX) — cut power across levels 1-4 for roughly 17.5 hours to full resolution. Onset was around 04:23-04:28 BST (Giganet observed feed loss at 04:23; BT's interruption began 04:28); Equinix declared all services resolved at 21:50 BST. Per The Register, the failure of an output static switch on a Galaxy UPS serving levels 1-4 tripped the facility fire alarm (no actual fire), causing a safety power-down of those floors that took out even the 'diverse' A and B feeds customers were contracted for. Customer evidence (Giganet) confirms both A and B feeds to one rack died simultaneously at ~04:23. Around 150 LINX members were affected, with named downstream ISPs — including BT, Sky and Virgin Media — knocked offline. Recovery was staggered and uneven: some services (e.g. BT) were back by ~08:03 while others waited into the evening, compounded because the site's access-control system also went dark, leaving customers queuing outside unable to badge in. Equinix's only official root-cause wording was 'a faulty UPS system'; it announced a review into power-supply resiliency but never published a formal RCA.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Power (2020-08-18)Trigger · Power2020-08-182020-08-18Primary fault at Equinix — LD8 (Docklands)Equinix LD8LD8 (Docklands)LD8Downstream service degraded by the fault: London Internet Exchange (LINX) peeringLondon Internet ExchangeDownstream service degraded by the fault: Customer colocation power across levels 1-4Customer colocation power acrossDownstream service degraded by the fault: Facility access-control / badge-in systemFacility access-control /badge-in…Downstream service degraded by the fault: Downstream ISP/telecom connectivityDownstream ISP/telecom

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Equinix
Data center
LD8 (Docklands)
Location
London, United Kingdom, Docklands
Date
2020-08-18

Impact & scale

Users affected
~150 LINX members plus 'hundreds of clients'; downstream ISPs/telcos named by DataCenterDynamics (Epsilon, SiPalto, EX Networks, Fast2Host, ICUK.net, Evoke Telecom) and by The Register (BT — interrupted ~04:28-08:03 BST — plus Sky, Virgin Media, and Exponential-e/M12-Giganet)
Financial
Not published — no financial, SLA-credit, or compensation figures disclosed by Equinix or affected parties
Scope
Major / severe (multi-floor de-energisation of a critical internet-exchange facility)
Services / systems down
  • London Internet Exchange (LINX) peering
  • Customer colocation power across levels 1-4
  • Facility access-control / badge-in system
  • Downstream ISP/telecom connectivity

Impact data & metrics

Onset (both A+B feeds lost)~04:23-04:24 BST, 18 Aug 2020
Fire alarm / detection time04:33 BST (per Equinix RFO)
Total end-to-end outage duration~17.5-18 h (~04:24 onset to 21:49 whole-site / ~22:20 final circuits)
LINX members directly affected~150
LINX PoPs impacted1 of 11 LON1/LON2 points of presence
Building levels lostLevels 1-4 of building 8/9
Busbar risers dropped2 (ISP8 and ISP9)
Fire suppression discharge0 (not activated; facility 'wasn't on fire')
BT Ethernet services restoredby 08:03 BST (interruption 04:28-08:03)
LINX power restoration start11:46 BST
All LINX devices restored13:42 BST (~9 h after onset)
Whole-site stable power restored21:49 BST
First Equinix public statement12:04 BST ('faulty UPS')

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 8Outage duration (0–10) — how long service was degraded/down. — scored 8/10.Blast 9Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 9/10.
Magnitude 7.8 = blast 9×0.35 + users 8×0.25 + financial 5×0.20 + duration 8×0.20 (sub-scores 0–10 · weighted composite)

Blast radius scored high because LD8 hosts LINX, a critical UK internet exchange, so failure rippled to ~150 members and multiple downstream carriers including BT, Sky and Virgin Media. Full-resolution duration ~17.5h is severe for a Tier-grade facility, though recovery was staggered (some services back within hours). Users score reflects ~150 LINX members plus 'hundreds of clients' and named ISPs. Financial score is a mid estimate: real economic harm is evident but no GBP/USD, SLA-credit, or compensation figures were ever published, so it cannot be quantified.

Sequence of events (SOE)

Phased sequence of events2020-08-18 ~04:23-04:24 BST · TRIGGER — Onset: Equinix LD8 (8/9 Harbour Exchange) loses power on the affected train; Giganet reports losing 'both A+B feeds' to a rack (~04:23 per Giganet, ~04:24 per Skipsey). Minor onset spread (~04:23-04:30) reflects detection vs LINX-logged times.TRIGGER2020-08-18 ~04:23-04:24 BS2020-08-18 04:33 · TRIGGER — The output static switch of the Galaxy UPS fails and emits smoke, triggering the building 8/9 fire-detection system; per the Equinix RFO the failed node was the 'main static transfer switch common output cabinet' (RFO single-source, not re-verifiable).TRIGGER2020-08-18 04:332020-08-18 04:33 · DETECTION — Fire alarm annunciates and Equinix facilities technicians respond immediately (per RFO time; onset itself was ~04:24).DETECTION2020-08-18 04:332020-08-18 ~04:33 · DETECTION — Responders find smoke but no open fire — combustion products from the faulting solid-state switch, not a sustained flame; Exponential-e 'confirmed the facility wasn't on fire.'DETECTION2020-08-18 ~04:32020-08-18 ~04:33 · CASCADE — The fault forces the UPS to shut down, causing (per RFO) a 'total loss of all customer power supplies supported by busbar risers ISP8 and ISP9' (riser names RFO single-source).CASCADE2020-08-18 ~04:32020-08-18 ~04:33 · IMPACT — Levels 1-4 of building 8/9 lose power; the outage takes down 1 of 11 LINX LON1/LON2 points of presence, directly affecting approximately 150 LINX members.IMPACT2020-08-18 ~04:32020-08-18 ~04:33 · MITIGATION — On the fire alarm, 'all systems within the vicinity of the fault – from levels one to four – were powered down' and the IBX was 'evacuated', delaying hands-on remediation.MITIGATION2020-08-18 ~04:32020-08-18 ~04:33 · MITIGATION — Fire suppression does not discharge and is not needed (per RFO, 'not activated nor required to do so at any point') — correct behaviour for electrical smoke without sustained fire; the facility 'wasn't on fire.'MITIGATION2020-08-18 ~04:32020-08-18 ~04:33 · CASCADE — Customer A/B redundancy is defeated because the affected feeds derived from the same UPS; Giganet 'suspect[s] a lack of resiliency' and notes all its kit is 'dual fed with diverse A+B power feeds' yet both were lost.CASCADE2020-08-18 ~04:32020-08-18 by 08:03 BST · RECOVERY — BT Ethernet services restored to affected BT Enterprise/Global customers (interruption 04:28-08:03) — an early, partial recovery ahead of the wider site.RECOVERY2020-08-18 by 08:03 BS2020-08-18 ~08:42 BST · MITIGATION — Equinix and contractors begin migrating customer power supplies onto newly commissioned infrastructure ('a decidedly non-live migrate' of systems 'due to be replaced') rather than restoring the faulted cabinet in place (08:42 exact time is Interstatus-only, flagged).MITIGATION2020-08-18 ~08:42 BS2020-08-18 11:46 BST · RECOVERY — 'Power is now being restored to the site and LINX equipment is in the process of being brought live again.'RECOVERY2020-08-18 11:46 BS2020-08-18 12:04 BST · MITIGATION — Equinix issues its first public statement: 'Equinix engineers have diagnosed the root cause of the issue as a faulty UPS ... system' — an under-detailed early attribution relative to the customer-distributed RFO.MITIGATION2020-08-18 12:04 BS2020-08-18 13:42 BST · RECOVERY — 'All of LINX devices have been restored' — roughly nine hours after onset; LINX rode through at network level with the other ten PoPs unaffected.RECOVERY2020-08-18 13:42 BS2020-08-18 21:49 BST · RESTORED — Equinix confirms complete and stable power was restored to the whole site — roughly a 17.5-hour end-to-end event from ~04:24 onset.RESTORED2020-08-18 21:49 BS2020-08-18 ~22:20 BST · RESTORED — Final circuits reported restored, ~18 hours after the outage began (Tech Monitor update, 19 Aug).RESTORED2020-08-18 ~22:20 BS

Root cause

Trigger. In the early hours of 18 August 2020 (customer-observed onset ~04:24 BST; Equinix's stated onset 04:40), the output static switch of a Schneider 'Galaxy' UPS system serving levels 1 through 4 of building 8/9 at Equinix's LD8 (Docklands, former Telecity) IBX failed. Equinix's own diagnosis, relayed via a customer IBX update quoted by The Register, named the root cause as 'a faulty UPS (uninterrupted power supply) system,' and stated that 'the fire alarm was triggered by the failure of output static switch from Galaxy UPS system supporting levels 1, 2, 3, 4' in building 8/9. The output static switch is the solid-state transfer element on the UPS output that selects between the inverter path and the bypass/mains path; its failure is the point at which conditioned power to the downstream distribution serving those four levels was compromised. Mechanism and escalation. The static-switch failure triggered the site fire alarm even though there was no actual fire. In accordance with site safety procedures, all systems in the vicinity of the fault across levels 1 to 4 were then powered down and the IBX was evacuated. The escalation path was therefore: electrical fault -> fire-alarm activation -> procedural mass power-down of an entire four-level block. The blast radius was driven as much by the safety response to the alarm as by the electrical fault itself. Why redundancy did not protect customers. Customers at LD8 were provisioned with dual 'A+B diverse' power feeds, the standard arrangement intended to survive loss of one supply path. This did not hold. Customer Matthew Skipsey reported that the alarm 'triggered the loss of both our A+B diverse power feeds to our main rack since 04:24,' and LINX reported that 'Simultaneously we lost power to our own A and B power feeds and subsequently our equipment within LD8.' Because the affected racks were fed from within the levels 1-4 block that was powered down, the 'diverse' feeds shared a common upstream dependency, so the loss was common-mode rather than independent. No official Equinix statement explains WHY both diverse feeds were lost together; the common-mode conclusion is an inference from converging customer reports plus the fact that the whole levels 1-4 zone was de-energised. (Note: a further specific timing and a 'clearly unacceptable / lack of resiliency' complaint have been attributed to customer Giganet, but that could not be verified against an accessible source and is excluded here.) Failover and suppression. There is no evidence that an alternative source (generator or alternate UPS string) automatically picked up the affected racks; the safety-driven power-down removed power from the zone, and recovery required manual work, including at least one operator (M12/Giganet) reporting its LD8 core router had been 'connected to a temporary power feed.' Fire suppression was not involved because there was no fire; the fire ALARM, not suppression, was the propagation mechanism. Downstream escalation and recovery. The outage propagated into the wider internet because LD8 'hosts one of the eleven points of presence for LINX LON1 and LON2 peering LANs'; 'approximately 150 LINX members will have been directly affected,' and a small number of BT Enterprise/Global Ethernet customers 'experienced an interruption to their Ethernet services between 04:28 and 08:03.' LINX reported power being restored to the site from 11:46 and 'all of LINX devices have been restored' by 13:42. A whole-site restoration time of 21:49 has been cited elsewhere but could not be verified in the accessible sources. Caveat on authority. Equinix publicly named the cause (faulty UPS) and the failing component (Galaxy UPS output static switch) via a brief customer update quoted by The Register, but no detailed engineering post-incident report was published. The deeper questions -- why the static switch failed, why its failure raised a fire alarm, and why the A+B diverse feeds were commonly dependent -- are not answered by any authoritative Equinix root-cause document in the available evidence. Official confirmation. This pass newly established a genuine primary official source: the LINX (London Internet Exchange) incident statement, a first-party record from an operator directly affected inside LD8. I fetched it directly on 2026-08-01 and verified all three quotes verbatim. It officially confirms the outage mechanism — loss of Equinix-supplied power at ~04:30 BST and the SIMULTANEOUS loss of both A and B feeds — and relays Equinix's confirmation of whole-site power restoration by 21:49 BST. Critical honesty note: LINX does NOT use the word 'UPS' and gives no root cause, so the widely-reported UPS attribution is corroborated by the mechanism but NOT confirmed by any retrieved primary Equinix document; Equinix's customer RFO remains non-public and unreachable here (press pages 403/404, archive.org blocked, WebSearch budget exhausted). officialPostmortem is set true on the strength of the LINX first-party official statement, with the UPS-specific cause flagged as press-sourced rather than officially confirmed.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What happened

In the early hours of 18 August 2020, the output static (transfer) switch of a Schneider Electric Galaxy UPS serving levels 1-4 of building 8/9 at Equinix's LD8 IBX (8/9 Harbour Exchange, Docklands) failed in service. It emitted smoke that tripped the fire-detection system at ~04:33 without a sustained fire, and the UPS shut down, dropping power to the affected floors. Because customers' A and B feeds both derived from this single UPS, dual-corded loads lost both cords together. LD8 hosts 1 of 11 LINX LON1/LON2 points of presence, so ~150 LINX members were directly affected, alongside ISPs including Giganet, Sky, Virgin Media and BT. Whole-site stable power was not restored until 21:49 — an end-to-end event of ~17.5-18 hours.

Why it was severe (the SPOF)

The severity came from topology, not just a failed UPS. Redundant A/B power only protects loads if the two paths are independent; at LD8 they converged on one UPS and its output static switch, so a single component fault defeated the diversity customers were paying for. Giganet lost 'both A+B feeds' despite 'diverse A+B power feeds', and an engineer relayed by Giganet's Matthew Skipsey described the rebuild as 'independent A+B PDUs and then route to separate UPS systems' where 'previously A+B went to a single UPS'. The remediation, by construction, confirms the pre-incident coupling.

Detection, safety and the long recovery

Detection and safety systems behaved correctly: the fire alarm annunciated immediately, fire suppression did not discharge (the site 'wasn't on fire'), and levels 1-4 were powered down and the IBX evacuated per procedure. But those correct safety actions, combined with a recovery path that required migrating customer load onto newly commissioned power infrastructure rather than resetting a switch — 'a decidedly non-live migrate' of gear 'due to be replaced' — drove the ~17.5-18-hour duration. Recovery was staged: BT services back by 08:03, LINX power restoration from 11:46, all LINX devices by 13:42, whole site by 21:49.

Evidence quality and open unknowns

The core facts are well-corroborated across three verified sources: LINX's official statement, Tech Monitor, and The Register (which independently quotes an Equinix update naming the Galaxy static-switch failure and levels 1-4). The most granular claims — the 'main static transfer switch common output cabinet' phrasing, the ISP8/ISP9 riser names, the exact 04:33 alarm time, and 'suppression not activated' — trace only to the Equinix RFO republished by Comtec, whose URL now 404s with no Wayback snapshot, so they are single-source and flagged. Undisclosed: the Galaxy model line, unit age, battery chemistry, the precise physical failure mode inside the static switch, and whether the SPOF was formally 'known' beforehand (that characterisation appears only in an unverified secondary NOC log).

Technical deep-dive

A UPS output static transfer switch is the solid-state (thyristor/SCR) element that transfers the critical output bus between the inverter source and the static-bypass (raw mains) source without a break; in Galaxy topologies a common output cabinet combines module outputs onto the load busbars. At LD8 this power train fed levels 1-4 of building 8/9 (verified via The Register's Equinix-update quote), and per the Equinix RFO the output distributed via two vertical busbar risers, ISP8 and ISP9 (RFO single-source, not re-verifiable). When the static switch faulted it did two things at once: it emitted enough smoke to annunciate the fire-detection system without an open flame (The Register: Exponential-e confirmed the site \"wasn't on fire\"), and it forced the UPS offline, collapsing the downstream supply. Because the A and B customer feeds both ultimately derived from this single UPS, the failure propagated past the redundancy boundary: dual-corded loads lost both feeds together — Giganet reported losing \"both A+B feeds\" at ~04:23-04:24 (Tech Monitor / The Register). This is the classic signature of a shared-component single point of failure in a nominally A/B-diverse system: diversity exists upstream and downstream but is nullified at the one place the paths physically merge. Detection worked (fire alarm ~04:33; technicians responded immediately per the RFO), and suppression correctly did not discharge because the event was electrical smoke rather than a sustained fire. The dominant contributor to the ~17.5-18-hour duration was therefore not detection latency but (a) a safety-driven evacuation and power-down of levels 1-4 (The Register: \"all systems within the vicinity of the fault – from levels one to four – were powered down\"), and (b) the need to migrate customer load onto newly commissioned power infrastructure rather than reset a switch in place — Giganet described Equinix \"migrating power supplies onto the new infrastructure following the earlier fault,\" and industry observer John Leach called it \"power systems that were due to be replaced [that] have failed early and are NOW being replaced. A decidedly non-live migrate\" (Tech Monitor). Recovery staged out: BT Ethernet services restored by 08:03 (The Register), LINX power restoration began 11:46, all LINX devices restored 13:42, and whole-site stable power at 21:49 (LINX), with final circuits reported ~22:20 (Tech Monitor).

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home