← All incidents
Incident dossier · Rank #25

Cloudflare Control-Plane & Analytics Outage — Flexential PDX-04 Power Failure (Nov 2023)

Cloudflare 2023-11-02 40h 41m core impact PowerHuman error

On 2 November 2023 a Portland General Electric (PGE) unplanned maintenance event dropped one of two independent utility feeds into Flexential's PDX-04 data center near Hillsboro, Oregon. Roughly three hours later a ground fault on a 12,470-volt PGE transformer triggered protective systems that, by design, shut down fast — but the same protective trip also shut down all ten of PDX-04's generators, which were unusually running in parallel with the surviving utility feed. UPS batteries meant to bridge to generator power began failing after only about four minutes, and the entire facility went dark. Cloudflare's control plane and analytics ran primarily across three Hillsboro-area data centers on an assumption that any one could fail, but two critical log-processing services (Kafka and ClickHouse) existed only in PDX-04 and had dependents in the supposedly highly-available cluster; PDX-04 also held more than a third of the HA cluster and Cloudflare's largest analytics cluster. That hidden single-site dependency — a failure mode Cloudflare had never tested by fully taking PDX-04 offline — broke the redundancy design. Cloudflare's non-automated failover to European disaster-recovery sites was slow, and cold-start recovery of servers and services ran into 4 November.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Power (2023-11-02)Trigger · Power2023-11-022023-11-02Primary fault at Cloudflare — Flexential PDX-04CloudflareFlexential PDX-04Flexential PDX-04Downstream service degraded by the fault: Cloudflare dashboard and API (configuration changes)Cloudflare dashboard and APIDownstream service degraded by the fault: AnalyticsAnalyticsDownstream service degraded by the fault: Raw log services (down full duration for most customers)Raw log servicesDownstream service degraded by the fault: Cloudflare StreamCloudflare Stream+2 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Cloudflare
Data center
Flexential PDX-04
Location
Hillsboro, United States, Portland/Hillsboro, Oregon
Date
2023-11-02

Impact & scale

Users affected
Cloudflare control-plane customers globally: dashboard and API configuration changes failed, analytics were unavailable, and raw log services were down for most customers for the entire incident. The data plane (edge proxy/CDN, WAF, DNS serving cached config) largely kept serving traffic. Exact customer counts are not stated in the primary postmortem; third-party figures (e.g. site/domain counts) are screening-only and were not re-verified.
Financial
Not disclosed in primary source
Scope
SEV-1 / Major (control-plane, not data-plane, total loss)
Services / systems down
  • Cloudflare dashboard and API (configuration changes)
  • Analytics
  • Raw log services (down full duration for most customers)
  • Cloudflare Stream
  • Magic WAN
  • Logpush (some logs dropped unrecoverably)

Impact data & metrics

Utility feed lost (precursor)1 of 2 independent PGE feeds dropped at 08:50 UTC, ~2h50m before the fault
Ground-fault voltage12,470 V (high-voltage transformer fault forcing fast protective shutdown)
UPS bridge performanceBatteries began failing after only ~4 minutes vs their intended ~10-minute bridge
Facility generation10 generators (inclusive of redundant units), rated to carry PDX-04 at full load — all taken offline by one protective trip
Total-power-loss window11:44–12:01 UTC (2 Nov) — every DC customer lost power
Generator restart delayGenerators restarted 12:48 UTC — ~1h04m after routers dropped / ~1h after UPS began failing
Operator notification lag~44 minutes (Cloudflare routers offline 11:44 → first Flexential notice 12:28 UTC)
Single-site blast radius>1/3 of the high-availability cluster machines + Cloudflare's largest analytics cluster housed in PDX-04
Hard single-site dependencyKafka + ClickHouse (log processing) existed ONLY in PDX-04, with dependents in the HA cluster
Total incident duration~40h41m (11:44 UTC Nov 2 → 04:25 UTC Nov 4); CEO characterized customer pain as 'the last 36 hours'
Raw log availabilityDown for most customers for the entire duration of the incident (total, not intermittent)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 7Users affected (0–10) — breadth of the user/customer population impacted. — scored 7/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 8Outage duration (0–10) — how long service was degraded/down. — scored 8/10.Blast 8Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 8/10.
Magnitude 7.2 = blast 8×0.35 + users 7×0.25 + financial 5×0.20 + duration 8×0.20 (sub-scores 0–10 · weighted composite)

Blast radius (8) and duration (8) dominate: a single-site power loss removed >1/3 of the high-availability cluster plus the largest analytics cluster and broke services meant to survive one-site loss, over a ~36-40 hour window running into a second calendar day. Users score (7) reflects a global control-plane/analytics/raw-log outage that nonetheless spared the data plane (traffic kept flowing). Financial (5) is a placeholder — no dollar impact is disclosed in the primary source; the reputational cost was high (public CEO apology). Scores are structural, not from a verified customer count, which the primary postmortem does not provide.

Sequence of events (SOE)

Phased sequence of events2023-11-02 08:50 UTC · TRIGGER — Portland General Electric (PGE) has an unplanned maintenance event affecting one of the two independent power feeds into Flexential PDX-04; the building loses one utility feed.TRIGGER2023-11-02 08:50 U2023-11-02 ~08:50 UTC · MITIGATION — Flexential powers up on-site generators to supplement the lost feed; facility now runs on generators plus the one remaining utility feed.MITIGATION2023-11-02 ~08:50 U2023-11-02 08:50-11:40 UTC · DETECTION — Notification/observability gap: Flexential does not tell Cloudflare it failed over to generators, and none of Cloudflare's observability tools detect the power-source change.DETECTION2023-11-02 08:50-11:40 U2023-11-02 ~11:40 UTC · CASCADE — A ground fault occurs on a PGE transformer at PDX-04 — believed (unconfirmed by Flexential/PGE) to be the step-down transformer for the second, still-running feed; high-voltage 12,470 V lines involved.CASCADE2023-11-02 ~11:40 U2023-11-02 ~11:40 UTC · CASCADE — Fast protective shutdown clears the fault but ALSO shuts down all of PDX-04's generators — a protection-coordination failure taking the remaining utility feed and on-site generation together.CASCADE2023-11-02 ~11:40 U2023-11-02 11:44 UTC · DETECTION — Cloudflare is first alerted to trouble only when the two routers connecting the facility to the world go offline — no prior notification from Flexential.DETECTION2023-11-02 11:44 U2023-11-02 ~11:44 UTC · IMPACT — UPS battery bank rated for ~10 minutes begins failing after only ~4 minutes; depletion inferred from Cloudflare's own equipment failing.IMPACT2023-11-02 ~11:44 U2023-11-02 ~11:44 UTC · CASCADE — Flexential's access-control system, not powered by the battery backups, goes offline — impeding physical entry for manual generator restart.CASCADE2023-11-02 ~11:44 U2023-11-02 11:44-12:01 UTC · IMPACT — With generators not restarted, the UPS batteries run out and ALL customers of the data center lose power (total facility de-energisation).IMPACT2023-11-02 11:44-12:01 U2023-11-02 ~12:01 UTC · CASCADE — Logical propagation: Kafka and ClickHouse — hosted only in PDX-04 — go down, taking dependent high-availability-cluster services (control plane + analytics/logging) with them.CASCADE2023-11-02 ~12:01 U2023-11-02 ~12:01 UTC · MITIGATION — Blast radius contained to the control/analytics plane: Cloudflare's edge network and security services keep serving traffic — the data plane is not impacted.MITIGATION2023-11-02 ~12:01 U2023-11-02 12:28 UTC · DETECTION — Flexential sends Cloudflare its first notification of the incident — about 44 minutes after Cloudflare's routers had already gone dark.DETECTION2023-11-02 12:28 U2023-11-02 (during outage) · MITIGATION — Manual generator restart is hampered by three factors reported by employees (unofficial): offline access control, difficulty, and thin staffing — 'security and an unaccompanied technician who had only been on the job for a week.'MITIGATION2023-11-02 (duri2023-11-02 12:48 UTC · RECOVERY — Flexential manually restarts generators, restoring partial power to the facility.RECOVERY2023-11-02 12:48 U2023-11-02 ~12:48 UTC · CASCADE — On attempting to re-energise Cloudflare's circuits, the circuit breakers are found faulty; cause undetermined (ground fault, surge, or pre-existing defect).CASCADE2023-11-02 ~12:48 U2023-11-02 (afternoon) · RECOVERY — Cloudflare works to fail control-plane and analytics services onto surviving infrastructure while breakers are replaced; data plane continues serving.RECOVERY2023-11-02 (afte2023-11-02 22:48 UTC · RESTORED — Flexential replaces the failed circuit breakers, restores both utility feeds, and confirms clean power to PDX-04 — about 11 hours after the ground fault.RESTORED2023-11-02 22:48 U2023-11-04 04:25 UTC · RESTORED — Cloudflare completes full restoration of control-plane and analytics services following clean-power restoration.RESTORED2023-11-04 04:25 U

Root cause

SPECIFIC INITIATING EVENT (an electrical de-energisation cascade — there was NO fire, no combustion, no suppression discharge): At approximately 11:40 UTC on 2 November 2023 a ground fault occurred on a Portland General Electric (PGE) transformer inside Flexential's PDX-04 building in Hillsboro, Oregon. Cloudflare states this was "very bad" because it involved "high voltage (12,470 volt) power lines," and that it was, on Cloudflare's own words, "the transformer that stepped down power from the grid for the second feed that was still running as it entered the data center" — though Cloudflare could not get confirmation from Flexential or PGE. No make, model, age, insulation chemistry, or manufacturer of that transformer is disclosed in any public source, and this uncertainty is explicit in the record. EXACT FAILURE MECHANISM: "Electrical systems are designed to quickly shut down to prevent damage" on a high-voltage ground fault; the protection operated, but the decisive perverse effect was that "the protective measure also shut down all of PDX-04's generators." A single utility-side transformer ground fault was thereby converted into simultaneous loss of BOTH the remaining utility feed and on-site generation. Framed as an engineering interpretation (not a Flexential admission), this reads as a protection-coordination / selectivity shortfall: the fault-isolation scheme was not selective enough to trip only the faulted section without also de-energising the standby generation meant to ride through exactly this event. LATENT / PRE-EXISTING CONDITIONS: (1) Degraded starting posture — at 08:50 UTC PGE had "an unplanned maintenance event affecting one of their independent power feeds," so PDX-04 had run on generators-plus-one-feed for ~2h50m before the ground fault struck the remaining feed's transformer; redundancy margin was already half-consumed. (2) UPS ride-through collapse — the battery bank "supposedly sufficient to power the facility for approximately 10 minutes" instead "started to fail after only 4 minutes" (Cloudflare inferred this from its own equipment failing); the ~4-minute reality versus ~10-minute rating is consistent with a degraded or under-capacity string, though Cloudflare does not itself diagnose the cause. (3) Faulty circuit breakers on Cloudflare's circuits, "discovered to be faulty" only on re-energisation, with Cloudflare openly stating "We don't know if the breakers failed due to the ground fault or some other surge as a result of the incident, or if they'd been bad before" — a possible latent equipment defect masked until the outage. MAINTENANCE / TESTING / STAFFING LATENT ROOT: Manual generator restart was hindered because "Flexential's access control system was not powered by the battery backups, so it was offline," physically impeding entry. Cloudflare explicitly flags the staffing account as unconfirmed — "While we haven't gotten official confirmation, we have been told by employees that three things hampered getting the generators back online" — one being that "the overnight shift consisted of security and an unaccompanied technician who had only been on the job for a week," i.e. no experienced operations or electrical expert on site. On Cloudflare's own side the latent root was a redundancy-testing gap: "we had never tested fully taking the entire PDX-04 facility offline," so hidden single-region dependencies (Kafka and ClickHouse existing only in PDX-04) inside the supposedly high-availability control plane went undetected, and products had been declared generally available without formally requiring migration to redundant backends — which Cloudflare labels plainly: "That was a mistake." Facility-side blame is Cloudflare's account; Flexential's and PGE's own maintenance/inspection records are not public and were not independently regulator-verified.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Not a fire — an electrical de-energisation cascade

No public source describes any combustion, fire-detection actuation, suppression discharge, evacuation, or fire-service response. The event is a power-loss cascade: a 12,470 V ground fault on a PGE transformer at Flexential PDX-04, protection that cleared the fault but 'also shut down all of PDX-04's generators,' and a UPS bank that failed in ~4 minutes against a ~10-minute rating. Treating this as a fire would misdiagnose the entire failure spine; the correct forensic frame is protection coordination, ride-through, and manual restart.

Redundancy consumed before the initiating fault

The facility was already degraded when the ground fault struck. PGE's 08:50 UTC unplanned maintenance had removed one of two independent feeds, and Flexential was supplementing with generators. For ~2h50m PDX-04 ran on generators plus a single feed — half its utility redundancy already gone. The ground fault then hit the transformer for that remaining feed, and the protective trip took the generators too. A latent maintenance window on the utility side thus set up the conditions for a single fault to become total loss.

Physical single-site fault, logical single-site dependency

The most consequential lesson is architectural. Physically, only PDX-04 went dark and Cloudflare's data plane kept serving traffic. But the control plane and analytics collapsed because Kafka and ClickHouse existed only in PDX-04 while high-availability-cluster services depended on them. A system marketed as highly available carried an untested single-region dependency. Cloudflare is candid that products reached general availability 'without formally requiring' redundant backends and that 'we had never tested fully taking the entire PDX-04 facility offline' — 'That was a mistake.'

Recovery friction: observability, communication, and staffing

Cloudflare had no signal of the power-source change and was first alerted only when its routers died at 11:44 UTC; Flexential's first notification came at 12:28 UTC. Manual generator restart — the only path, since automatic restart did not occur — was slowed by an access-control system off the battery backup and, per unconfirmed employee accounts, an overnight shift of 'security and an unaccompanied technician who had only been on the job for a week.' Partial power returned at 12:48 UTC but Cloudflare's breakers were found faulty; clean power was not confirmed until 22:48 UTC, roughly 11 hours after the fault.

Evidence quality and disclosed gaps

Every load-bearing claim traces to Cloudflare's own official post-mortem, an authoritative vendor account of its own outage. Its limits are disclosed and preserved here: Cloudflare could not get Flexential/PGE confirmation on the transformer, does not know whether the breakers were pre-faulted, inferred UPS depletion from its own failing gear, and flags the staffing detail as unconfirmed employee testimony. No independent regulatory, court, or PGE/Flexential engineering record is public, so facility-side findings rest on one party's narrative and should not be read as regulator-verified.

Technical deep-dive

This incident is an electrical power-loss cascade, not a fire — there was no combustion, no fire-detection actuation, no clean-agent or water-based suppression discharge, no evacuation, and no fire-service response in any source. The correct forensic spine is the de-energisation and protection sequence. POWER ARCHITECTURE AS DESIGNED: PDX-04 was fed by two independent PGE utility feeds, backed by on-site diesel generators and a UPS battery bank sized for roughly 10 minutes of ride-through — a standard utility/generator/UPS topology intended to survive loss of a single feed. Cloudflare leases space here where it houses "our largest analytics cluster as well as more than a third of the machines for our high availability cluster," making PDX-04 one of three Portland-area facilities and its most consequential for the control plane. DEFENCE-IN-DEPTH DEFEATED LAYER BY LAYER: Layer 1 (utility redundancy) was half-defeated at 08:50 UTC when PGE maintenance removed one feed; Flexential correctly started generators to supplement — but "counter to best practices, Flexential did not inform Cloudflare that they had failed over to generator power," and none of Cloudflare's observability tools detected the power-source change, so situational awareness was blind. Layer 2 (fault protection vs. standby-generation coordination) failed at ~11:40 UTC: the ground fault on the 12,470 V transformer tripped protection that "also shut down all of PDX-04's generators" — electrically correct, but mis-coordinated, taking the generators with it. Layer 3 (UPS ride-through) failed fast: batteries rated ~10 minutes collapsed in ~4 minutes, leaving no time for manual generator restart. Between 11:44 and 12:01 UTC the UPS batteries ran out and all customers of the data center lost power. DETECTION / SENSING: Grid protection relays detected the fault; Cloudflare detected nothing until its own hardware died — "We were first notified of issues in the data center when the two routers that connect the facility to the rest of the world went offline at 11:44 UTC." Flexential's first notification to Cloudflare came only at 12:28 UTC, ~44 minutes later. The UPS batteries effectively became the last sensor, with depletion inferred from Cloudflare's own equipment failing. RECOVERY OBSTACLES: Manual generator restart was the only path (automatic restart did not occur) and was slowed by an offline access-control system (not on battery backup) and thin, inexperienced overnight staffing. Generators were manually restarted around 12:48 UTC for partial power, but Cloudflare's circuit breakers "were discovered to be faulty," cause undetermined. Full clean power came only at 22:48 UTC when "Flexential replaced our failed circuit breakers, restored both utility feeds, and confirmed clean power" — about 11 hours after the ground fault. Full Cloudflare service restoration completed by 04:25 UTC on 4 November. LOGICAL BLAST RADIUS: Physically the loss was contained to PDX-04, but it propagated into Cloudflare's logical redundancy because "Kafka and ClickHouse — were only available in PDX-04 but had services that depended on them" running in the high-availability cluster. Crucially the data plane held: Cloudflare's network and security services continued to serve traffic; only the control plane (dashboard/API) and analytics/logging failed. A single-site physical fault exposed a single-site logical dependency hiding inside a system marketed as highly available.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home