← All incidents
Incident dossier · Rank #34

Fastly Global CDN Outage — Latent Software Bug (June 8, 2021)

Fastly 2021-06-08 1h 0m core impact SoftwareNetwork

An undiscovered software bug introduced in a mid-May Fastly software deployment lay dormant until a single customer configuration change triggered it on 8 June 2021, disabling roughly 85% of Fastly's global CDN and taking major sites (Amazon, Reddit, gov.uk, The New York Times, Twitch) offline worldwide for about an hour.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Software (2021-06-08)Trigger · Software2021-06-082021-06-08Primary fault at Fastly — Fastly global edge CDN networkFastlyFastly global edge CDN networkFastly global edge CDN networkDownstream service degraded by the fault: ~85% of Fastly CDN globally~85% of Fastly CDN globallyDownstream service degraded by the fault: Major customer sites (Amazon, Reddit, gov.uk, NYT, Twitch)Major customer sites

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Fastly
Data center
Fastly global edge CDN network
Location
Global, Global
Date
2021-06-08

Impact & scale

Users affected
Global users of major sites (Amazon, Reddit, gov.uk, NYT, Twitch, Spotify, Shopify)
Financial
Not published
Scope
Sev-1 global CDN (software)
Services / systems down
  • ~85% of Fastly CDN globally
  • Major customer sites (Amazon, Reddit, gov.uk, NYT, Twitch)

Impact data & metrics

Share of global network returning errors~85%
Network operating normally within 49 minutes95%
Latent-defect dwell time (deploy to trigger)~27 days (May 12 to June 8)
Detection latency (onset to monitoring alert)~1 minute (09:47 to 09:48 UTC)
Onset to public status post~11 minutes (09:47 to 09:58 UTC)
Onset to root cause identified~40 minutes (09:47 to 10:27 UTC)
Onset to services begin recovering / 95% normal~49 minutes (09:47 to 10:36 UTC)
Onset to majority of services recovered~73 minutes (09:47 to 11:00 UTC)
Onset to incident fully mitigated~2h48m (09:47 to 12:35 UTC)
Onset to permanent bug-fix deployment start~7h38m (09:47 to 17:25 UTC)
Confirmed major sites taken offline (verbatim subset)6 (Amazon, Reddit, gov.uk, NYT, Guardian, BBC)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 7Users affected (0–10) — breadth of the user/customer population impacted. — scored 7/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 3Outage duration (0–10) — how long service was degraded/down. — scored 3/10.Blast 8Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 8/10.
Magnitude 6.2 = blast 8×0.35 + users 7×0.25 + financial 5×0.20 + duration 3×0.20 (sub-scores 0–10 · weighted composite)

A single customer config change triggered a latent bug that disabled ~85% of Fastly's CDN, taking major sites offline globally for ~1h — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events2021-05-12 · TRIGGER — Fastly begins a software deployment that introduces the latent bug 'that could be triggered by a specific customer configuration under specific circumstances'; the defect lies dormant across the global edge fleet for ~27 days, having passed pre-deployment QA and testing.TRIGGER2021-05-122021-06-08 09:47 UTC · TRIGGER — Onset — a customer pushes a valid configuration change meeting the 'specific circumstances,' activating the dormant bug (the ignition analogue: a legitimate action strikes the latent defect).TRIGGER2021-06-08 09:47 U2021-06-08 09:47 UTC · CASCADE — ~85% of Fastly's global edge network begins returning errors instead of serving content; because the defective code was already resident fleet-wide, the fault propagates network-wide almost immediately — the software analogue of near-instant global spread.CASCADE2021-06-08 09:47 U2021-06-08 09:48 UTC · DETECTION — Fastly monitoring identifies the global disruption 'within one minute' (the fire-detection-system analogue) — no smoke/heat sensors; observability tooling flags the network-wide error surge.DETECTION2021-06-08 09:48 U2021-06-08 09:48 UTC · IMPACT — Major downstream sites go dark globally — Reddit, gov.uk, Amazon, The New York Times, The Guardian and the BBC 'become unavailable' (press also cited Twitch, Spotify, GitHub and others).IMPACT2021-06-08 09:48 U2021-06-08 09:58 UTC · DETECTION — Fastly publishes a public status post acknowledging the incident (~11 min after onset) — the external-alarm analogue; self-reported, no emergency-services call.DETECTION2021-06-08 09:58 U2021-06-08 09:48–10:27 UTC · MITIGATION — Fastly internal incident-response engineering (the 'emergency services' analogue) works the network-wide error condition between detection and root-cause identification; no external brigade dispatched — self-detected and self-driven.MITIGATION2021-06-08 09:48–10:27 U2021-06-08 10:27 UTC · MITIGATION — Fastly Engineering identifies the customer configuration as the cause (~40 min after onset) — the root-cause-identification step preceding suppression — then moves to isolate and disable it.MITIGATION2021-06-08 10:27 U2021-06-08 10:36 UTC · MITIGATION — Impacted services begin to recover as the triggering configuration is disabled and healthy state is forced across the fleet (the manual 'suppression' analogue — no automatic suppression existed to halt the fault).MITIGATION2021-06-08 10:36 U2021-06-08 10:36 UTC · RECOVERY — Per Fastly's own summary, 'Within 49 minutes, 95% of our network was operating as normal' — restoration by config-disable, not de-energisation: edge PoPs stayed powered and were returned to the healthy request path (the electrical-isolation analogue was logical, not physical).RECOVERY2021-06-08 10:36 U2021-06-08 11:00 UTC · RECOVERY — Majority of services recovered — global end-user disruption lasted roughly 49 minutes for most affected sites and up to ~73 minutes for the tail.RECOVERY2021-06-08 11:00 U2021-06-08 12:35 UTC · RECOVERY — Incident declared mitigated (~2h48m after onset) — containment complete, error rates normalized across the network.RECOVERY2021-06-08 12:35 U2021-06-08 12:44 UTC · RECOVERY — Fastly's public status post is marked resolved — end of the customer-facing incident window.RECOVERY2021-06-08 12:44 U2021-06-08 17:25 UTC · RESTORED — Deployment of the permanent bug fix begins across the network — 'We created a permanent fix for the bug and began deploying it at 17:25' — the true permanent 'suppression,' separate from the earlier config-disable mitigation.RESTORED2021-06-08 17:25 U2021-06-08 (post-incident) · RESTORED — Fastly apologizes — 'This outage was broad and severe, and we're truly sorry for the impact to our customers' — and commits to investigating why QA/testing missed the bug and to 'evaluate ways to improve our remediation time.'RESTORED2021-06-08 (post

Root cause

This was a global software control-plane failure of Fastly's edge CDN, with no physical ignition source, equipment, or chemistry; mapped to the incident schema, the "ignition source" analogue is a single defective code change. Per Fastly's official post-mortem (Nick Rockwell, fastly.com/blog/summary-of-june-8-outage): "On May 12, we began a software deployment that introduced a bug that could be triggered by a specific customer configuration under specific circumstances." That deployment is the proximate defect origin. The bug lay dormant — latent across essentially the entire global edge fleet — for roughly 27 days before any input exercised it. The failure-mechanism analogue is a latent, globally-distributed logic defect gated on an input condition: it required a particular customer configuration state to activate, so until that state existed the defective code path was never reached and the fault was invisible in normal operation. Fastly never publicly disclosed the specific module, function, service, or code path; the post-mortem names only "a bug" triggered by "a specific customer configuration under specific circumstances." Fastly's edge is Varnish/VCL-derived, but the official record does not name the language or component, so the precise mechanism is marked disclosed-unknown. The trigger — the match-strike analogue — was a legitimate customer action, not an error: "a customer pushed a valid configuration change that included the specific circumstances that triggered the bug, which caused 85% of our network to return errors." Fastly stressed the change was "valid," locating fault in its own software rather than the customer. Neither the customer nor the specific config was ever identified. The latent root cause — the design/inspection-lapse analogue — is a gap in Fastly's software quality-assurance and deploy-validation regime: the defect passed pre-deployment QA on May 12 and survived nearly a month undetected in production. Fastly conceded it would "figure out why we didn't detect the bug during our software quality assurance and testing processes." A second latent root is the absence of a rollout/blast-radius guardrail: a single valid config edit propagated the error condition to ~85% of the global network almost simultaneously, with no canary, staged rollout, or automated circuit-breaker catching it before worldwide impact — Fastly separately committed to "evaluate ways to improve our remediation time." No physical maintenance, inspection, detection, or suppression hardware is implicated because none was involved.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Nature of the incident: a software control-plane failure, not a physical event

There was no fire, facility damage, detection or suppression hardware, evacuation, or emergency-services response. Mapped to the incident-forensic schema, this was a global failure of Fastly's edge CDN control plane in which the 'ignition source' analogue is a single defective code change. Every schema element (ignition, detection, suppression, isolation, evacuation) is expressed here as its software analogue, grounded verbatim in Fastly's official post-mortem and Wikipedia; where an analogue does not exist (physical de-energisation, brigade dispatch), it is marked not-applicable rather than fabricated.

The latent defect and its ~27-day dwell

The proximate defect origin is the May 12, 2021 deployment that 'introduced a bug that could be triggered by a specific customer configuration under specific circumstances.' The defect was distributed fleet-wide but dormant, gated on an input state, so it passed pre-deployment QA and survived nearly a month in production undetected. On June 8 at 09:47 UTC a customer pushed a 'valid configuration change' meeting those circumstances, activating the bug and causing ~85% of the network to return errors. Fastly never disclosed the specific module, function, language, or code path — that remains disclosed-unknown.

Detection strong, containment weak

Automated monitoring flagged the disruption 'within one minute' (onset 09:47, alert 09:48 UTC), a public status post followed at 09:58, root cause was identified at 10:27, and per Fastly '95% of our network was operating as normal' within 49 minutes (~10:36 UTC), with the majority recovered by ~11:00 and full mitigation at 12:35. The lesson is that fast detection did not prevent near-instant global impact: absent canary/staged rollout or an automated circuit-breaker, a single valid config propagated the fault to ~85% of the fleet before any human could intervene.

Downstream blast radius and concentration risk

Wikipedia confirms Reddit, gov.uk, Amazon, The New York Times, The Guardian and the BBC 'become unavailable'; press cited many more (Twitch, Spotify, GitHub, Stripe, PayPal and others), grounded only where verbatim. One provider-level control-plane fault simultaneously darkened otherwise unrelated global services — a concrete demonstration of single-CDN concentration risk. Market-impact and share-price figures reported by press are excluded as unsourced against the official record.

What Fastly committed to fix

Fastly deployed a permanent bug fix beginning 17:25 UTC the same day, apologized ('This outage was broad and severe, and we're truly sorry for the impact to our customers'), and committed to investigating why QA/testing missed the bug and to 'evaluate ways to improve our remediation time.' The unstated-but-clear design lesson — blast-radius guardrails so one valid config edit cannot reach ~85% of the network — is flagged as inferred rather than an explicit Fastly commitment.

Technical deep-dive

This was a software/CDN control-plane failure with no fire, facility, detection hardware, suppression system, evacuation, or emergency-services response; the incident-forensic schema is mapped to its software analogues, each grounded in Fastly's official post-mortem (fastly.com/blog/summary-of-june-8-outage) and Wikipedia. Latent defect and trigger: A May 12, 2021 software deployment "introduced a bug that could be triggered by a specific customer configuration under specific circumstances." The defect was globally distributed but dormant, gated on a particular configuration state. On June 8, "a customer pushed a valid configuration change that included the specific circumstances that triggered the bug, which caused 85% of our network to return errors." Because the defective code was already resident on nearly the entire edge fleet, the triggering config propagated the error condition network-wide almost immediately — the software analogue of near-instantaneous global spread rather than a localized, propagating burn. Blast radius: ~85% of Fastly's global network returned errors instead of serving content. (Contemporaneous press widely reported HTTP 503 responses in the request path; the specific status code is press-attributed and is NOT stated in Fastly's post-mortem, so it is not asserted here as official fact.) Detection (= fire-detection system analogue): automated monitoring/observability, not smoke or heat sensors. "We detected the disruption within one minute" — onset 09:47 UTC, monitoring flagged the global disruption 09:48 UTC. A public status post followed at 09:58 UTC. Detection was fast and effective; the gaps were upstream (pre-production QA never caught the latent defect) and downstream (no automated containment before global impact). Suppression / mitigation (= suppression system analogue): no automatic suppression halted the fault. Mitigation was manual engineering intervention — Fastly Engineering "identified the customer configuration" at 10:27 UTC (~40 min after onset), then isolated and disabled it. Recovery followed fast: per Fastly's own summary, "Within 49 minutes, 95% of our network was operating as normal" (by ~10:36 UTC, where impacted services "began to recover"); the majority of services had recovered by ~11:00 UTC. The incident was declared mitigated at 12:35 UTC and the status post resolved at 12:44 UTC. The permanent "suppression" — the actual bug fix — was a separate, later action: "We created a permanent fix for the bug and began deploying it at 17:25." Isolation / de-energisation (= electrical-isolation analogue): not applicable as physical de-energisation. Functionally, ~85% of Fastly's edge PoPs stayed powered and running but failed the request path; recovery came by disabling the offending configuration and forcing healthy state, not by cutting power. Evacuation and emergency services: not applicable — no personnel or physical site were endangered; "impact" was to traffic and customers, not people, and there was no fire brigade or 911/999 dispatch. The analogue is Fastly's internal incident-response engineering, which self-detected (no external call), engaged within one minute, and drove isolation → recovery → mitigation → permanent fix across the 09:47–17:25 UTC window. Downstream impact: per Wikipedia, the outage caused "Reddit, gov.uk, and Amazon, along with major news sources such as The New York Times, The Guardian, and the BBC, to become unavailable." Press cited many more sites (Twitch, PayPal, Spotify, GitHub, Stripe, HBO Max, CNN, FT, Pinterest, Kickstarter, Vimeo, Shopify, Hulu), but only the Wikipedia-confirmed subset is grounded verbatim. Economic/market-impact figures and Fastly's share-price move that day were reported by press but not quantified in the official post-mortem, and are excluded as unsourced-official. Fastly closed: "This outage was broad and severe, and we're truly sorry for the impact to our customers," and committed to investigating the QA/testing gap and improving remediation time.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home