Equinix LD8 Docklands UPS Failure: ~17.5-Hour Power Outage Disrupts LINX and ~150 Members
On 18 August 2020 a faulty UPS system at Equinix's LD8 Docklands data centre — home to the London Internet Exchange (LINX) — cut power across levels 1-4 for roughly 17.5 hours to full resolution. Onset was around 04:23-04:28 BST (Giganet observed feed loss at 04:23; BT's interruption began 04:28); Equinix declared all services resolved at 21:50 BST. Per The Register, the failure of an output static switch on a Galaxy UPS serving levels 1-4 tripped the facility fire alarm (no actual fire), causing a safety power-down of those floors that took out even the 'diverse' A and B feeds customers were contracted for. Customer evidence (Giganet) confirms both A and B feeds to one rack died simultaneously at ~04:23. Around 150 LINX members were affected, with named downstream ISPs — including BT, Sky and Virgin Media — knocked offline. Recovery was staggered and uneven: some services (e.g. BT) were back by ~08:03 while others waited into the evening, compounded because the site's access-control system also went dark, leaving customers queuing outside unable to badge in. Equinix's only official root-cause wording was 'a faulty UPS system'; it announced a review into power-supply resiliency but never published a formal RCA.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Equinix
- Data center
- LD8 (Docklands)
- Location
- London, United Kingdom, Docklands
- Date
- 2020-08-18
Impact & scale
- Users affected
- ~150 LINX members plus 'hundreds of clients'; downstream ISPs/telcos named by DataCenterDynamics (Epsilon, SiPalto, EX Networks, Fast2Host, ICUK.net, Evoke Telecom) and by The Register (BT — interrupted ~04:28-08:03 BST — plus Sky, Virgin Media, and Exponential-e/M12-Giganet)
- Financial
- Not published — no financial, SLA-credit, or compensation figures disclosed by Equinix or affected parties
- Scope
- Major / severe (multi-floor de-energisation of a critical internet-exchange facility)
- London Internet Exchange (LINX) peering
- Customer colocation power across levels 1-4
- Facility access-control / badge-in system
- Downstream ISP/telecom connectivity
Impact data & metrics
| Onset (both A+B feeds lost) | ~04:23-04:24 BST, 18 Aug 2020 |
| Fire alarm / detection time | 04:33 BST (per Equinix RFO) |
| Total end-to-end outage duration | ~17.5-18 h (~04:24 onset to 21:49 whole-site / ~22:20 final circuits) |
| LINX members directly affected | ~150 |
| LINX PoPs impacted | 1 of 11 LON1/LON2 points of presence |
| Building levels lost | Levels 1-4 of building 8/9 |
| Busbar risers dropped | 2 (ISP8 and ISP9) |
| Fire suppression discharge | 0 (not activated; facility 'wasn't on fire') |
| BT Ethernet services restored | by 08:03 BST (interruption 04:28-08:03) |
| LINX power restoration start | 11:46 BST |
| All LINX devices restored | 13:42 BST (~9 h after onset) |
| Whole-site stable power restored | 21:49 BST |
| First Equinix public statement | 12:04 BST ('faulty UPS') |
Magnitude profile
Blast radius scored high because LD8 hosts LINX, a critical UK internet exchange, so failure rippled to ~150 members and multiple downstream carriers including BT, Sky and Virgin Media. Full-resolution duration ~17.5h is severe for a Tier-grade facility, though recovery was staggered (some services back within hours). Users score reflects ~150 LINX members plus 'hundreds of clients' and named ISPs. Financial score is a mid estimate: real economic harm is evident but no GBP/USD, SLA-credit, or compensation figures were ever published, so it cannot be quantified.
Sequence of events (SOE)
- TRIGGER Onset: Equinix LD8 (8/9 Harbour Exchange) loses power on the affected train; Giganet reports losing 'both A+B feeds' to a rack (~04:23 per Giganet, ~04:24 per Skipsey). Minor onset spread (~04:23-04:30) reflects detection vs LINX-logged times.
- TRIGGER The output static switch of the Galaxy UPS fails and emits smoke, triggering the building 8/9 fire-detection system; per the Equinix RFO the failed node was the 'main static transfer switch common output cabinet' (RFO single-source, not re-verifiable).
- DETECTION Fire alarm annunciates and Equinix facilities technicians respond immediately (per RFO time; onset itself was ~04:24).
- DETECTION Responders find smoke but no open fire — combustion products from the faulting solid-state switch, not a sustained flame; Exponential-e 'confirmed the facility wasn't on fire.'
- CASCADE The fault forces the UPS to shut down, causing (per RFO) a 'total loss of all customer power supplies supported by busbar risers ISP8 and ISP9' (riser names RFO single-source).
- IMPACT Levels 1-4 of building 8/9 lose power; the outage takes down 1 of 11 LINX LON1/LON2 points of presence, directly affecting approximately 150 LINX members.
- MITIGATION On the fire alarm, 'all systems within the vicinity of the fault – from levels one to four – were powered down' and the IBX was 'evacuated', delaying hands-on remediation.
- MITIGATION Fire suppression does not discharge and is not needed (per RFO, 'not activated nor required to do so at any point') — correct behaviour for electrical smoke without sustained fire; the facility 'wasn't on fire.'
- CASCADE Customer A/B redundancy is defeated because the affected feeds derived from the same UPS; Giganet 'suspect[s] a lack of resiliency' and notes all its kit is 'dual fed with diverse A+B power feeds' yet both were lost.
- RECOVERY BT Ethernet services restored to affected BT Enterprise/Global customers (interruption 04:28-08:03) — an early, partial recovery ahead of the wider site.
- MITIGATION Equinix and contractors begin migrating customer power supplies onto newly commissioned infrastructure ('a decidedly non-live migrate' of systems 'due to be replaced') rather than restoring the faulted cabinet in place (08:42 exact time is Interstatus-only, flagged).
- RECOVERY 'Power is now being restored to the site and LINX equipment is in the process of being brought live again.'
- MITIGATION Equinix issues its first public statement: 'Equinix engineers have diagnosed the root cause of the issue as a faulty UPS ... system' — an under-detailed early attribution relative to the customer-distributed RFO.
- RECOVERY 'All of LINX devices have been restored' — roughly nine hours after onset; LINX rode through at network level with the other ten PoPs unaffected.
- RESTORED Equinix confirms complete and stable power was restored to the whole site — roughly a 17.5-hour end-to-end event from ~04:24 onset.
- RESTORED Final circuits reported restored, ~18 hours after the outage began (Tech Monitor update, 19 Aug).
Root cause
Contributing factors
- Design SPOF: the A and B customer feeds both derived from a single Schneider Galaxy UPS, so one output-static-switch failure collapsed the affected supply and defeated A/B power diversity — Giganet lost 'both A+B feeds' at ~04:23-04:24 (Tech Monitor / The Register).
- Coupled A/B power paths at the UPS: the post-incident fix uses 'independent A+B PDUs and then route to separate UPS systems' whereas 'previously A+B went to a single UPS' (engineer relayed by Giganet's Matthew Skipsey, via Tech Monitor) — confirming diversity was effectively 1N at the shared UPS/static switch.
- Solid-state static-switch vulnerability with no evidenced condition-monitoring intercept: the thyristor/SCR output switch failed in service and only became visible once it emitted smoke and tripped the fire alarm; no verified source evidences predictive/thermal monitoring that caught degradation earlier (physical failure mode undisclosed).
- Safety-driven power-down and evacuation delayed hands-on remediation: 'all systems within the vicinity of the fault – from levels one to four – were powered down' and the IBX was 'evacuated' (The Register) — correct safety behaviour that nonetheless added time before engineers could work the fault.
- Recovery required migrating load onto newly commissioned power infrastructure rather than a fast switch reset — a 'decidedly non-live migrate' of power systems 'due to be replaced' (Tech Monitor) — stretching restoration across ~17.5-18 hours.
- Concentration of critical peering infrastructure: LD8 hosted 1 of 11 LINX LON1/LON2 points of presence, so a single power-domain failure directly affected approximately 150 LINX members (LINX / The Register).
- Communication gap during the event: Equinix's first public statement came at 12:04 (Tech Monitor) and customers (Skipsey) called the lack of communication 'abysmal', while access-control systems were knocked offline, forcing manual two-way-radio operation.
Correction of errors (COE)
- Re-architect A/B feeds via independent A+B PDUs routing to separate UPS systems, decoupling the shared output path so no single UPS/static switch feeds both cords
- Audit all LD8 distribution boards, busbar risers (ISP8/ISP9) and common-output cabinets for hidden A/B convergence nodes and eliminate shared single points
- Deploy condition-monitoring (thermal/harmonic trending) on solid-state static switches to catch degradation before in-service failure
- Establish zoned fault-isolation procedures allowing work on a faulted UPS cabinet without full-facility evacuation/power-down halting recovery
- Provide affected customers a timely, complete, verifiable RFO and keep access-control systems on protected power
Lessons learnt
- Nominal A/B redundancy is only as strong as its convergence points: a single shared UPS/output static switch nullified dual-corded diversity for affected LD8 customers — Giganet lost 'both A+B feeds' despite 'diverse A+B power feeds' — so redundancy must be verified physically end-to-end, not assumed from the one-line diagram.
- A 'faulty UPS' headline can obscure the real lesson: the trigger was one solid-state output-transfer component (a 'failing output static switch in the Galaxy UPS system (sold by Schneider)', Tech Monitor), and the damage came from where that component sat in the topology, not merely that a UPS failed.
- Fire suppression not firing is not a suppression failure: for electrical smoke without a sustained flame the correct behaviour is non-discharge — the facility 'wasn't on fire' (The Register) — so post-incident review should not chase a phantom suppression fault.
- Recovery time is dominated by the safe path, not detection: detection was immediate (~04:33) but restoration ran ~17.5-18 h because levels 1-4 were powered down and load had to be migrated onto newly commissioned infrastructure ('a decidedly non-live migrate', Tech Monitor) — resilience planning must include a fast, rehearsed alternate-energisation path.
- Concentration risk is real for shared internet infrastructure: a single Equinix power domain directly affected ~150 LINX members even though only 1 of 11 PoPs was hit (LINX) — critical members should dual-site across facilities.
- Keep distinct incidents distinct: The Register flags a separate July 2016 Telecity outage that knocked ~10% of BT subscribers offline; that older event must not be conflated with this 2020 Galaxy static-switch failure.
Improvements & remediation
- Audit every 'diverse' A/B feed pair end-to-end to confirm true physical and electrical separation with no shared UPS, output static switch, or busway single point of failure.
- Provide genuinely independent redundant UPS paths (separate systems, separate output switching) for critical loads, verified by pull-the-plug testing.
- Isolate or independently power life-safety interlocks so a fire-alarm/EPO trip cannot needlessly de-energise redundant power sides across whole floors.
- Put physical access-control on protected/backup power (or an offline-capable fallback) so customers can reach their equipment during an outage.
- Adopt a formal incident-communications standard: acknowledge within a defined window and provide rolling restoration ETAs.
- Publish a formal RCA/COE after major incidents to enable customer trust and industry learning.
- Institute periodic 2N-failover drills and independent resiliency audits rather than relying on design intent.
Comprehensive analysis
What happened
In the early hours of 18 August 2020, the output static (transfer) switch of a Schneider Electric Galaxy UPS serving levels 1-4 of building 8/9 at Equinix's LD8 IBX (8/9 Harbour Exchange, Docklands) failed in service. It emitted smoke that tripped the fire-detection system at ~04:33 without a sustained fire, and the UPS shut down, dropping power to the affected floors. Because customers' A and B feeds both derived from this single UPS, dual-corded loads lost both cords together. LD8 hosts 1 of 11 LINX LON1/LON2 points of presence, so ~150 LINX members were directly affected, alongside ISPs including Giganet, Sky, Virgin Media and BT. Whole-site stable power was not restored until 21:49 — an end-to-end event of ~17.5-18 hours.
Why it was severe (the SPOF)
The severity came from topology, not just a failed UPS. Redundant A/B power only protects loads if the two paths are independent; at LD8 they converged on one UPS and its output static switch, so a single component fault defeated the diversity customers were paying for. Giganet lost 'both A+B feeds' despite 'diverse A+B power feeds', and an engineer relayed by Giganet's Matthew Skipsey described the rebuild as 'independent A+B PDUs and then route to separate UPS systems' where 'previously A+B went to a single UPS'. The remediation, by construction, confirms the pre-incident coupling.
Detection, safety and the long recovery
Detection and safety systems behaved correctly: the fire alarm annunciated immediately, fire suppression did not discharge (the site 'wasn't on fire'), and levels 1-4 were powered down and the IBX evacuated per procedure. But those correct safety actions, combined with a recovery path that required migrating customer load onto newly commissioned power infrastructure rather than resetting a switch — 'a decidedly non-live migrate' of gear 'due to be replaced' — drove the ~17.5-18-hour duration. Recovery was staged: BT services back by 08:03, LINX power restoration from 11:46, all LINX devices by 13:42, whole site by 21:49.
Evidence quality and open unknowns
The core facts are well-corroborated across three verified sources: LINX's official statement, Tech Monitor, and The Register (which independently quotes an Equinix update naming the Galaxy static-switch failure and levels 1-4). The most granular claims — the 'main static transfer switch common output cabinet' phrasing, the ISP8/ISP9 riser names, the exact 04:33 alarm time, and 'suppression not activated' — trace only to the Equinix RFO republished by Comtec, whose URL now 404s with no Wayback snapshot, so they are single-source and flagged. Undisclosed: the Galaxy model line, unit age, battery chemistry, the precise physical failure mode inside the static switch, and whether the SPOF was formally 'known' beforehand (that characterisation appears only in an unverified secondary NOC log).
Technical deep-dive
References & provenance
- official-postmortem Outage: Equinix LD8 Reachability — LINX Technical Blog“At approximately 04:30BST Equinix lost power to their LD8 datacentre”https://www.linx.net/outage-equinix-ld8-reachability/
- vendor-status LINX Incidents Log - LD8 power outage 18 August 2020“Simultaneously we lost power to our own A and B power feeds and subsequently our equipment within LD8.”https://www.linx.net/incidents-log/
- press Tech Monitor (CBR): Equinix Outage: Major Power Failure Leaves Data Centre Customers Fuming“Various customers identified the issue as a failing output static switch in the Galaxy UPS system (sold by Schneider).”https://www.techmonitor.ai/hardware/data-centres/equinix-outage
- news Equinix LD8 data center experiences major outage“Equinix engineers have diagnosed the root cause of the issue as a faulty UPS (uninterrupted power supply) system.”https://www.datacenterdynamics.com/en/news/equinix-ld8-data-center-experiences-major-outage/
- news Equinix LD8 data center experiences major outage (Giganet incident report)“We lost both A+B feeds to 1 of our 2 Equinix LD8 racks at approximately 4.23am ... This follows a UPS failure, which then triggered the fire alarm in the data center according to reports from Equinix.”https://www.datacenterdynamics.com/en/news/equinix-ld8-data-center-experiences-major-outage/
- news Equinix confirms review under way into UPS failure at Docklands datacentre“Equinix is conducting a review into the resiliency of its power supplies after an outage at one of its London-based datacentres blighted the operations of hundreds of clients on 18 August 2020.”https://www.computerweekly.com/news/252487861/Equinix-confirms-review-under-way-into-UPS-failure-at-Docklands-datacentre
- news Outage: Faulty UPS at data centre housing London Internet Exchange causes grief for ISPs and telcos alike“the fire alarm was triggered by the failure of output static switch from Galaxy UPS system supporting levels 1, 2, 3, 4 ... 150 of their members are affected by this outage.”https://www.theregister.com/2020/08/18/outage_london_internet_exchange/
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.