← All incidents
Incident dossier · Rank #4

Meta Global BGP + DNS Outage — Backbone Self-Disconnection Takes Facebook, Instagram & WhatsApp Offline Worldwide

Meta (Facebook) 2021-10-04 7h 11m core impact NetworkHuman error

On 4 October 2021 a single command issued during routine backbone maintenance — intended only to assess global backbone capacity — unintentionally took down every connection in Facebook's backbone, completely severing all of its data centers from the internet. With the backbone gone, Facebook's DNS-serving locations followed their designed safety behavior and declared themselves unhealthy, withdrawing the BGP route advertisements for the authoritative nameservers. That withdrawal pulled Facebook's DNS off the global internet: recursive resolvers everywhere lost any path to the nameservers, so Facebook, Messenger, Instagram, WhatsApp, Mapillary and Oculus all failed simultaneously rather than degrading gracefully. Recovery was slow because primary and out-of-band network access were also down, forcing engineers to travel on-site; the intentionally hardened physical and system security of the facilities made restarting systems difficult, with recovery advancing after a team reached and reset servers at the Santa Clara, California data center. The outage ran roughly six to seven hours (15:39 UTC start to general restoration around 22:50 UTC).

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2021-10-04)Trigger · Network2021-10-042021-10-04Primary fault at Meta (Facebook) — Facebook Global Backbone / Authoritative DNS (Santa Clara data center recovery site)MetaFacebook Global Backbone / Authoritative DNS (Santa Clara data center recovery site)Facebook Global Backbone /Downstream service degraded by the fault: FacebookFacebookDownstream service degraded by the fault: MessengerMessengerDownstream service degraded by the fault: InstagramInstagramDownstream service degraded by the fault: WhatsAppWhatsApp+4 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Meta (Facebook)
Data center
Facebook Global Backbone / Authoritative DNS (Santa Clara data center recovery site)
Location
Menlo Park, USA, California
Date
2021-10-04

Impact & scale

Users affected
All users of Facebook, Messenger, Instagram, WhatsApp, Mapillary and Oculus globally (family of apps served the same withdrawn authoritative DNS); precise user count not verified in primary sources
Financial
Not verified in primary Meta or Cloudflare sources; widely-cited revenue/market-cap figures remain unsourced and are excluded here
Scope
SEV-1 / Global total outage (characterized by CNBC as Facebook's worst outage since 2008)
Services / systems down
  • Facebook
  • Messenger
  • Instagram
  • WhatsApp
  • Mapillary
  • Oculus
  • Internal engineering tooling
  • Physical badge / building access

Impact data & metrics

Outage start15:39 UTC, 4 Oct 2021
BGP routing restored~21:50 UTC
DNS services resumed22:05 UTC
General platform restoration~22:50 UTC
Total duration~6-7 hours (431 min, 15:39->22:50 UTC)
External DNS-availability window (Cloudflare vantage)~5.5 hours (~15:50 to ~21:20 UTC)
Meta properties globally unavailable6 (Facebook, Messenger, Instagram, WhatsApp, Mapillary, Oculus)
Backbone connectivity lost100% — complete data-center-to-internet disconnection
Historical scaleWorst Facebook outage since 2008 (per CNBC, via Wikipedia)
Published numbered COE action items0 (no enumerated corrective list published)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 10Users affected (0–10) — breadth of the user/customer population impacted. — scored 10/10.Financial 8Financial impact (0–10) — direct + consequential cost. — scored 8/10.Duration 7Outage duration (0–10) — how long service was degraded/down. — scored 7/10.Blast 10Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 10/10.
Magnitude 9.0 = blast 10×0.35 + users 10×0.25 + financial 8×0.20 + duration 7×0.20 (sub-scores 0–10 · weighted composite)

Blast radius and user impact maxed: every Meta property globally offline at once plus loss of internal tooling and badge access. Duration ~6-7h is severe but bounded. Financial score is a qualitative estimate — CNBC called it the worst outage since 2008 — because specific dollar figures are not verified in the primary sources.

Sequence of events (SOE)

Phased sequence of events2021-10-04 ~15:40 UTC · TRIGGER — During a routine backbone maintenance job a capacity-assessment command is issued that unintentionally withdraws all backbone connections between Meta data centers; Cloudflare simultaneously sees a peak of BGP routing changes leaving Facebook's network — 'That's when the trouble began.'TRIGGER2021-10-04 ~15:40 U2021-10-04 ~15:40 UTC · CASCADE — The change-audit tool that should have blocked the command fails due to a bug: 'a bug in that audit tool prevented it from properly stopping the command.'CASCADE2021-10-04 ~15:40 U2021-10-04 ~15:40 UTC · CASCADE — The entire backbone is removed from operation, effectively disconnecting Meta's data centers globally.CASCADE2021-10-04 ~15:40 U2021-10-04 ~15:40 UTC · CASCADE — By health-check design, DNS servers cannot reach the data centers, declare themselves unhealthy, and withdraw their BGP advertisements — a self-protective mechanism firing on a global false-positive.CASCADE2021-10-04 ~15:40 U2021-10-04 ~15:40 UTC · DETECTION — External detection is near-immediate: Cloudflare observes the peak of routing changes as backbone connections begin dropping.DETECTION2021-10-04 ~15:40 U2021-10-04 ~15:50 UTC · IMPACT — facebook.com ceases to resolve on public resolvers as authoritative nameserver prefixes become unreachable — Cloudflare's 1.1.1.1 records it 'stopped being available at around 15:50 UTC.'IMPACT2021-10-04 ~15:50 U2021-10-04 15:58 UTC · CASCADE — Facebook stops announcing routes to its DNS prefixes via BGP — 'Facebook had stopped announcing the routes to their DNS prefixes,' effectively disconnecting itself from the Internet.CASCADE2021-10-04 15:58 U2021-10-04 ~15:58 UTC · CASCADE — Public resolvers 1.1.1.1, 8.8.8.8 and other major resolvers begin 'issuing (and caching) SERVFAIL responses,' propagating the failure across the DNS ecosystem.CASCADE2021-10-04 ~15:58 U2021-10-04 (during outage) · IMPACT — A global retry storm forms as clients worldwide re-query relentlessly: 'DNS resolvers worldwide handling 30x more queries than usual' — collateral load on third-party infrastructure.IMPACT2021-10-04 (duri2021-10-04 (during outage) · DETECTION — Internal detection is crippled: 'The total loss of DNS broke many of the internal tools we'd normally use to investigate and resolve outages like this.'DETECTION2021-10-04 (duri2021-10-04 (during outage) · MITIGATION — Remote data-center access is impossible because the networks are down; secure physical-access protocols must be activated — 'it took extra time to activate the secure access protocols needed to get people onsite.'MITIGATION2021-10-04 (duri2021-10-04 (during outage) · MITIGATION — Engineers are physically dispatched into hardened data centers to debug and restart the systems, slowed by hardware 'designed to be difficult to modify even when you have physical access to them.'MITIGATION2021-10-04 (duri2021-10-04 ~21:00 UTC · RECOVERY — Cloudflare sees renewed BGP activity from Facebook's network (peaking 21:17 UTC) as backbone connectivity is restored on-site.RECOVERY2021-10-04 ~21:00 U2021-10-04 ~21:00 UTC · RECOVERY — Power-up is staged cautiously because data centers had power dips 'in the range of tens of megawatts' and a sudden reversal 'could put everything from electrical systems to caches at risk.'RECOVERY2021-10-04 ~21:00 U2021-10-04 ~21:20 UTC · RESTORED — Facebook DNS returns to availability on Cloudflare's 1.1.1.1 ('returned at 21:20 UTC'); services progressively recover as backbone connectivity is restored, with full reconnection noted by 21:28 UTC.RESTORED2021-10-04 ~21:20 U2021-10-05 · RESTORED — Meta (Santosh Janardhan) publishes the detailed post-incident writeup: capacity-assessment command + audit-tool bug + DNS self-withdrawal + tooling/physical-access difficulties, following the initial 10-04 statement.RESTORED2021-10-05

Root cause

There was no ignition source, combustion, or equipment fire — this was a pure software-and-procedure failure, so the conventional root-cause chain (make/model/age/chemistry of a failed component) is inapplicable by the nature of the incident, and the record is explicit that router/DNS-server hardware makes, models and firmware were never causal. Meta states the event was "caused not by malicious activity, but an error of our own making" (engineering.fb.com/2021/10/05). Proximate trigger — a single erroneous maintenance command. During a routine backbone maintenance job, "a command was issued with the intention to assess the availability of global backbone capacity, which unintentionally took down all the connections in our backbone network, effectively disconnecting Meta's data centers globally" (engineering.fb.com/2021/10/05). Latent root — a defective safety guardrail. The change-validation tooling that exists specifically to catch this class of error was itself broken: "Our systems are designed to audit commands like these to prevent mistakes like this, but a bug in that audit tool prevented it from properly stopping the command" (engineering.fb.com/2021/10/05). The failure is twofold: an erroneous command reached production, and the automated guardrail meant to stop it had failed. Design flaw that converted a backbone fault into total global disconnection — an automated self-protective mechanism firing on a correlated, global false-positive. "Our DNS servers disable those BGP advertisements if they themselves can not speak to our data centers, since this is an indication of an unhealthy network connection. In the recent outage the entire backbone was removed from operation, making these locations declare themselves unhealthy and withdraw those BGP advertisements" (engineering.fb.com/2021/10/05). The health-check worked exactly as designed, but because the backbone loss made every DNS location simultaneously "unhealthy," authoritative DNS withdrew itself from the Internet worldwide — the safety mechanism amplified rather than contained the failure. Recovery-compounding latent conditions — the loss of DNS also disabled Meta's own diagnostic tooling ("The total loss of DNS broke many of the internal tools we'd normally use to investigate and resolve outages like this"), and physical/system security hardening slowed on-site manual recovery ("once you're inside, the hardware and routers are designed to be difficult to modify even when you have physical access to them"; "it took extra time to activate the secure access protocols needed to get people onsite") (engineering.fb.com/2021/10/05).

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Incident Overview

On 4 October 2021, Facebook, Instagram, WhatsApp, Messenger and Oculus went offline worldwide for roughly six hours. The cause was not a fire, hardware failure, or attack but a procedural error compounded by a broken guardrail. During routine backbone maintenance, a command intended only to assess global backbone capacity instead withdrew every backbone connection between Meta's data centers, and the audit tool designed to reject such a command had a bug that let it through. Meta characterizes the event as 'caused not by malicious activity, but an error of our own making' (engineering.fb.com/2021/10/05).

The Three-Layer Cascade: Backbone to BGP to DNS

The blast radius grew because three layers were tightly coupled. Losing the backbone (layer 1) severed every DNS point-of-presence from the data centers. Meta's DNS servers are designed to withdraw their BGP advertisements (layer 2) when they cannot reach the data centers, on the assumption that unreachability means a locally unhealthy site. Because the whole backbone vanished at once, every DNS location judged itself unhealthy simultaneously and withdrew in unison — removing facebook.com's authoritative nameservers (layer 3) from the global routing table. Cloudflare externally clocked the sequence: routing-change peak ~15:40 UTC, DNS unavailable ~15:50 UTC, DNS-prefix withdrawal 15:58 UTC (blog.cloudflare.com/october-2021-facebook-outage).

Why the Safety Mechanisms Amplified the Failure

Two protective systems turned a recoverable fault into a total, self-inflicted disconnection. First, the change-audit tool — a safety interlock meant to stop dangerous commands — was itself defective, so the erroneous command executed unchecked. Second, the DNS health-check acted precisely as designed, but its scope was unbounded: a mechanism built to route around one bad site instead removed authoritative DNS from the entire Internet when the failure was global. When DNS fell, public resolvers (1.1.1.1, 8.8.8.8 and others) began issuing and caching SERVFAIL, and worldwide retries drove resolver load to about 30x normal — spreading collateral impact onto infrastructure Meta does not own.

Recovery Constraints and the Physical-Access Bottleneck

Recovery was slow for reasons that were themselves latent design conditions. The remote management path relied on the same network that was down, and the DNS loss disabled many of Meta's internal investigation tools. Engineers therefore had to be physically dispatched into data centers hardened against intruders, where activating secure access protocols took extra time and where hardware is deliberately 'difficult to modify even when you have physical access.' Restoration was then intentionally staged, because data centers showed power dips of tens of megawatts and a sudden reversal 'could put everything from electrical systems to caches at risk.' Cloudflare observed renewed BGP from ~21:00 UTC, DNS back by ~21:20 UTC, and full reconnection by 21:28 UTC.

Systemic Lessons for Correlated-Failure Design

The incident is a canonical study in correlated failure and self-amplifying safety systems. Guardrails must be tested as rigorously as the systems they protect; health-driven automation must be bounded so a global fault cannot trigger universal self-withdrawal; diagnostic and recovery tooling must not depend on the very service that can fail; and physical security must be balanced against a rehearsed break-glass recovery path. Above all, resilience planning must account for blast radius beyond one's own estate — here a single command rippled out to third-party resolvers across the entire Internet.

Technical deep-dive

The failure propagated across three coupled layers: backbone (physical/IP connectivity between data centers), BGP (Internet route announcements), and DNS (authoritative name resolution). During a routine backbone maintenance job, a capacity-assessment command withdrew all backbone connections at once (engineering.fb.com/2021/10/05: "unintentionally took down all the connections in our backbone network"). The audit tool that should have blocked the command failed due to a bug ("a bug in that audit tool prevented it from properly stopping the command"). With the backbone gone, every Meta DNS point of presence lost its path to Meta's data centers. By design, a Meta DNS server withdraws its BGP route advertisements when it cannot reach the data centers, treating unreachability as a signal of an unhealthy connection. Normally this reroutes traffic around one bad location; here every location was simultaneously cut off, so all of them withdrew their BGP advertisements at once. Cloudflare observed this externally: at ~15:40 UTC "a peak of routing changes from Facebook. That's when the trouble began," and at 15:58 UTC "Facebook had stopped announcing the routes to their DNS prefixes." With those withdrawals, Facebook had effectively disconnected itself from the Internet. Because facebook.com's authoritative nameservers were now unreachable at the routing layer, no resolver on Earth could complete a lookup; Cloudflare's 1.1.1.1 recorded facebook.com resolution as unavailable "at around 15:50 UTC." Public recursive resolvers began issuing and caching failures: "1.1.1.1, 8.8.8.8, and other major public DNS resolvers started issuing (and caching) SERVFAIL responses." Every app, browser and retrying client worldwide then hammered resolvers on a loop, producing a retry storm — Cloudflare reported "DNS resolvers worldwide handling 30x more queries than usual," a collateral load spike on infrastructure Meta does not own. Recovery could not use the normal remote path because the network carrying it was down, and DNS loss had also disabled Meta's internal diagnostic and access tooling. Engineers were physically dispatched into hardened data centers, where "it took extra time to activate the secure access protocols needed to get people onsite and able to work on the servers," further slowed by hardware "designed to be difficult to modify even when you have physical access." Restoration was deliberately staged: data centers were reporting "dips in power usage in the range of tens of megawatts," and "suddenly reversing such a dip in power consumption could put everything from electrical systems to caches at risk," so the backbone was brought back in controlled steps. Cloudflare saw renewed BGP activity from ~21:00 UTC (peaking 21:17 UTC), DNS availability return by ~21:20 UTC, and full reconnection by 21:28 UTC.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home