← All incidents
Incident dossier · Rank #32

Cloudflare June 2022 Outage — Config Change Withdraws 19 Core Data Centers

Cloudflare 2022-06-21 1h 15m core impact NetworkSoftware

A network-configuration change during a resilience upgrade to Cloudflare's Multi-Colo PoP architecture withdrew BGP routes at the company's 19 busiest core data centers, dropping roughly half of all Cloudflare requests globally for about 75 minutes.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2022-06-21)Trigger · Network2022-06-212022-06-21Primary fault at Cloudflare — Cloudflare 19 busiest core data centers (MCP)CloudflareCloudflare 19 busiest core data centers (MCP)Cloudflare 19 busiest core dataDownstream service degraded by the fault: Cloudflare CDN/proxy at 19 core locationsCloudflare CDN/proxy at 19 coreDownstream service degraded by the fault: Dependent customer sitesDependent customer sites

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Cloudflare
Data center
Cloudflare 19 busiest core data centers (MCP)
Location
Global, Global
Date
2022-06-21

Impact & scale

Users affected
~50% of Cloudflare requests worldwide; large share of major internet sites
Financial
Not published
Scope
Sev-1 global edge (config)
Services / systems down
  • Cloudflare CDN/proxy at 19 core locations
  • Dependent customer sites

Impact data & metrics

Core data centers taken offline19 (MCP-architecture spine locations)
Affected sites as share of total network~4% (Cloudflare's phrasing: 'only 4% of our total network')
Global request volume affected~50% of total requests impacted
Detection-to-incident-declaration~5 min (06:27 onset to 06:32 declaration)
Onset-to-root-cause-found~31 min (06:27 to 06:58)
Core recovery duration~75 min (06:27 onset to 07:42 final revert)
Total incident duration (to closure)~1h33m (06:27 to 08:00)
Automated rollback available at time of incident0 (no commit-confirm; recovery fully manual)
Benign pre-MCP deploy stages before impact2 (03:56 UTC and 06:17 UTC, zero impact)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 7Users affected (0–10) — breadth of the user/customer population impacted. — scored 7/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 4Outage duration (0–10) — how long service was degraded/down. — scored 4/10.Blast 8Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 8/10.
Magnitude 6.3 = blast 8×0.35 + users 7×0.25 + financial 5×0.20 + duration 4×0.20 (sub-scores 0–10 · weighted composite)

~50% of global requests dropped for ~75 min from a routing/config change at the 19 busiest core DCs — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events2022-06-21 03:56 UTC · TRIGGER — Change deployed to the first location. No impact — this site runs the older, non-MCP architecture, confirming the edit is inert outside the MCP spine tier.TRIGGER2022-06-21 03:56 U2022-06-21 06:17 UTC · TRIGGER — Change deployed to the busiest, non-MCP locations. Still no impact — the destructive path has not been touched.TRIGGER2022-06-21 06:17 U2022-06-21 06:27 UTC · TRIGGER — Rollout reaches MCP-enabled locations and the reordered AGGREGATES-OUT policy commits to the spines: this is when the incident started. Site-local terms now sit below REJECT-THE-REST and match 'then reject;' first.TRIGGER2022-06-21 06:27 U2022-06-21 06:27 UTC · IMPACT — All 19 MCP data centers taken offline — the change withdrew the site-local prefixes across the MCP spines (no sub-staggering within the MCP tier).IMPACT2022-06-21 06:27 U2022-06-21 06:27 UTC · CASCADE — Withdrawal of site-local prefixes severs direct reachability to the 19 locations AND the servers' reachability to origin servers: it 'removed our direct access to all the impacted locations, as well as removing the ability of our servers to reach origin servers.'CASCADE2022-06-21 06:27 U2022-06-21 06:27 UTC · CASCADE — Internal load balancer Multimog collapses: it 'could no longer forward requests between the servers' in the MCPs, so even surviving servers inside a data center cannot be reached east-west.CASCADE2022-06-21 06:27 U2022-06-21 06:27 UTC · IMPACT — Half of global request volume can no longer be served: the affected locations are 'only 4% of our total network' yet 'the outage impacted 50% of total requests.'IMPACT2022-06-21 06:27 U2022-06-21 06:32 UTC · DETECTION — Internal Cloudflare incident formally declared — ~5 minutes after onset. (Source does not name a specific monitoring alarm; detection was via the deploy's own impact and internal incident processes.)DETECTION2022-06-21 06:32 U2022-06-21 06:51 UTC · MITIGATION — First change made on a router to verify the root cause — engineers begin confirming the policy-ordering hypothesis on a live device.MITIGATION2022-06-21 06:51 U2022-06-21 06:58 UTC · DETECTION — Root cause found — the term-reordering of AGGREGATES-OUT is confirmed as the cause; ~31 minutes from onset. Revert work begins.DETECTION2022-06-21 06:58 U2022-06-21 06:58 UTC · MITIGATION — Reversion of the problematic change begins. Because direct access was withdrawn, engineers fall back to backup take-control procedures — no automated commit-confirm rollback exists.MITIGATION2022-06-21 06:58 U2022-06-21 07:00-07:40 UTC · CASCADE — Recovery self-inflicts delay: 'network engineers walked over each other's changes, reverting the previous reverts,' causing the problem to re-appear intermittently during manual reversion.CASCADE2022-06-21 07:00-07:40 U2022-06-21 07:42 UTC · RECOVERY — Final reverts completed; core routing restored ~75 minutes after the 06:27 onset. AGGREGATES-OUT is back to correct term ordering and site-locals re-advertised.RECOVERY2022-06-21 07:42 U2022-06-21 08:00 UTC · RESTORED — Incident closed — ~1h33m from onset at 06:27; core recovery achieved by 07:42 (~75 min).RESTORED2022-06-21 08:00 U

Root cause

SPECIFIC TRIGGER (ignition equivalent): The outage was ignited by a single BGP route-export-policy change pushed to the spine routers of Cloudflare's 19 MCP (Multi-Colo PoP) data centers during a resilience upgrade. The change reordered the terms of the export policy named AGGREGATES-OUT, relocating the two site-local advertisement terms — 4-ADV-SITE-LOCALS and 6-ADV-SITE-LOCALS — from the top of the policy to the bottom, i.e. beneath the catch-all term REJECT-THE-REST { then reject; }. Cloudflare states plainly: "the 4-ADV-SITE-LOCALS and 6-ADV-SITE-LOCALS terms moved from the top to the bottom." FAILURE MECHANISM: BGP export policies are evaluated first-match, top-to-bottom, and the first matching term is final. With the two site-local terms now sitting below REJECT-THE-REST, every site-local prefix matched "then reject;" first and was discarded before it could be advertised. The spine routers therefore withdrew the site-local prefixes for the affected MCP locations the instant the change committed: "we immediately stopped advertising our site-local prefixes." Per Cloudflare, this "removed our direct access to all the impacted locations, as well as removing the ability of our servers to reach origin servers." Two failures happened in the same stroke — the network lost its route INTO those data centers, and the servers inside them lost their route OUT to origins. DEVICE (make/model equivalent): The affected equipment is the MCP "spine" router tier forming a Clos-network mesh — "an added layer of routing that creates a mesh of connections." Cloudflare does not name the vendor. The post-mortem's configuration dialect (policy-statement terms, "then reject", "community add", and the recommended "commit-confirm" rollback) is consistent with Juniper Junos syntax, so the devices are INFERRED to be Juniper routers running Junos — this is an inference from the config vocabulary, NOT a stated fact in the source. There is no age/chemistry analogue; the salient hazardous property is architectural: the newer, centralized MCP tier ran a single shared export policy across the affected sites' spines, so one bad edit propagated identically. The companion change made on the older-architecture routers (adding BGP communities STATIC-ROUTE, SITE-LOCAL-ROUTE, TLL01, EUROPE) was harmless; the destructive effect came solely from the term-reordering on the MCP spines. LATENT ROOT (design / redundancy / change-management lapse): The proximate config error was allowed to reach global blast radius by latent weaknesses Cloudflare itself admits. (1) Fragile policy design — AGGREGATES-OUT depended on implicit term ordering plus an explicit REJECT-THE-REST catch-all, so a mere reorder silently inverted its behavior; the policy statement "will be redesigned to prevent an unintentional incorrect ordering." (2) Inadequate stagger scoping — although the change used a stagger procedure, "the stagger policy did not include an MCP data center until the final step," so no MCP spine was exercised until the change hit the MCP tier. (3) Steps too coarse and review blind to the hazard — "the steps weren't small enough to catch the error before it hit all of our spines." (4) No automated rollback safety-net — there was no "commit-confirm" auto-revert; its absence forced slow manual recovery and, per Cloudflare, "would have greatly reduced the Time-to-Resolve." The root cause is therefore not merely "an engineer reordered a policy" but a change-management and architecture design in which a single unvalidated edit to a shared policy could withdraw the busiest half of the network with no automatic backstop.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What actually broke

A BGP export-policy change on the spine routers of Cloudflare's 19 MCP data centers reordered policy terms in AGGREGATES-OUT, moving the site-local advertisement terms (4-ADV-SITE-LOCALS, 6-ADV-SITE-LOCALS) below the catch-all REJECT-THE-REST. Because BGP policies match first-term-wins top-to-bottom, the site-local prefixes matched 'then reject;' first and were withdrawn. That single stroke removed both inbound reachability to the affected data centers and the servers' outbound reachability to origins, and additionally broke the internal Multimog load balancer so requests could not be forwarded between servers inside an MCP.

Why it spread so far so fast

The impact was bounded by architecture, not geography: only MCP-converted sites failed, proven by the two earlier non-MCP deploy stages (03:56 and 06:17 UTC) causing zero impact. But the MCP tier ran a shared export policy, so once the rollout reached it at 06:27 UTC the fault applied across all 19 spines effectively at once. Those sites were 'only 4% of our total network' yet the outage 'impacted 50% of total requests' — a concentration of traffic that turned a single-tier config error into a global event.

Why recovery took ~75 minutes

Detection was fast (incident declared 06:32, ~5 min after onset) and root cause was found by 06:58 (~31 min). But recovery was slow because the failure removed engineers' own direct access to the sites they had to fix, forcing fallback to backup take-control procedures, and because there was no automated commit-confirm rollback. Manual reversion then regressed itself — 'network engineers walked over each other's changes, reverting the previous reverts' — pushing final restoration to 07:42 and closure to 08:00 UTC.

Systemic lessons and remediation

Cloudflare's own post-mortem attributes the blast radius to a procedural gap (stagger policy excluded MCP until the final step; steps too coarse to catch the error) and fragile policy design (order-dependent advertisement plus a catch-all reject). Committed fixes span Process (MCP-specific test/deploy procedures), Architecture (redesign the policy statement to prevent unintentional incorrect ordering), and Automation (improved stagger that exercises an MCP site early, plus automated commit-confirm rollback that 'would have greatly reduced the Time-to-Resolve'). The broader lesson is that a resilience upgrade is itself a change whose recovery path must not depend on the reachability it can remove.

Evidence quality and open gaps

All load-bearing facts are drawn from Cloudflare's first-party detailed post-mortem, which names the policy, the terms, the timeline, the metrics, and the remediation. The router vendor is not stated by Cloudflare; the Juniper/Junos identification is an inference from the config vocabulary (policy-statement terms, 'then reject', 'commit-confirm') and is disclosed as such. The specific monitoring signal that first flagged impact is not named in the source; detection is described via the deploy's own impact and internal incident processes.

Technical deep-dive

Cloudflare had been migrating its busiest data centers to a Multi-Colo PoP (MCP) architecture — a Clos network described as "an added layer of routing that creates a mesh of connections," intended to be a more flexible and resilient architecture. Ironically the outage occurred during work on that very architecture. On 21 June 2022 a network change was rolled out that, among benign edits, altered the BGP export policy AGGREGATES-OUT on the MCP spine routers. BGP export policies in this Junos-style syntax are ordered lists of terms evaluated top-to-bottom with first-match-wins semantics. AGGREGATES-OUT contained terms that advertised the site-local prefixes (4-ADV-SITE-LOCALS for IPv4, 6-ADV-SITE-LOCALS for IPv6) and a catch-all term REJECT-THE-REST { then reject; } intended to drop everything else. The change moved the two advertisement terms below REJECT-THE-REST. Because first match is final, the site-local prefixes now matched "then reject;" before ever reaching their advertisement terms, and the spines withdrew those prefixes. Two consequences fired at once. Externally, withdrawing the site-locals "removed our direct access to all the impacted locations, as well as removing the ability of our servers to reach origin servers" — north-south connectivity collapsed. Internally, the removal of these site-local prefixes also caused Cloudflare's internal load-balancing system Multimog "to stop working, as it could no longer forward requests between the servers" in the MCPs — east-west connectivity inside each MCP collapsed too. So even reaching a surviving server inside a data center could not help, because Multimog could no longer distribute the request. The blast radius was defined by architecture, not geography: only MCP-converted sites were affected. The earlier deploy stages at 03:56 UTC (first location) and 06:17 UTC (non-MCP locations) caused zero impact — evidence the change was destructive only on the MCP spines. Yet those 19 sites, though "only 4% of our total network," carried the traffic such that "the outage impacted 50% of total requests," so half of global request volume could no longer be served. Recovery was crippled by the outage's own mechanism. Having withdrawn the site-locals, engineers had lost direct access to the exact locations they needed to fix, and there was no automated commit-confirm rollback to self-heal. Manual reversion then collided with itself: "network engineers walked over each other's changes, reverting the previous reverts," which is why the final reverts slipped to 07:42 UTC despite the root cause being found at 06:58 UTC.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home