Cloudflare Control-Plane & Analytics Outage — Flexential PDX-04 Power Failure (Nov 2023)
On 2 November 2023 a Portland General Electric (PGE) unplanned maintenance event dropped one of two independent utility feeds into Flexential's PDX-04 data center near Hillsboro, Oregon. Roughly three hours later a ground fault on a 12,470-volt PGE transformer triggered protective systems that, by design, shut down fast — but the same protective trip also shut down all ten of PDX-04's generators, which were unusually running in parallel with the surviving utility feed. UPS batteries meant to bridge to generator power began failing after only about four minutes, and the entire facility went dark. Cloudflare's control plane and analytics ran primarily across three Hillsboro-area data centers on an assumption that any one could fail, but two critical log-processing services (Kafka and ClickHouse) existed only in PDX-04 and had dependents in the supposedly highly-available cluster; PDX-04 also held more than a third of the HA cluster and Cloudflare's largest analytics cluster. That hidden single-site dependency — a failure mode Cloudflare had never tested by fully taking PDX-04 offline — broke the redundancy design. Cloudflare's non-automated failover to European disaster-recovery sites was slow, and cold-start recovery of servers and services ran into 4 November.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Cloudflare
- Data center
- Flexential PDX-04
- Location
- Hillsboro, United States, Portland/Hillsboro, Oregon
- Date
- 2023-11-02
Impact & scale
- Users affected
- Cloudflare control-plane customers globally: dashboard and API configuration changes failed, analytics were unavailable, and raw log services were down for most customers for the entire incident. The data plane (edge proxy/CDN, WAF, DNS serving cached config) largely kept serving traffic. Exact customer counts are not stated in the primary postmortem; third-party figures (e.g. site/domain counts) are screening-only and were not re-verified.
- Financial
- Not disclosed in primary source
- Scope
- SEV-1 / Major (control-plane, not data-plane, total loss)
- Cloudflare dashboard and API (configuration changes)
- Analytics
- Raw log services (down full duration for most customers)
- Cloudflare Stream
- Magic WAN
- Logpush (some logs dropped unrecoverably)
Impact data & metrics
| Utility feed lost (precursor) | 1 of 2 independent PGE feeds dropped at 08:50 UTC, ~2h50m before the fault |
| Ground-fault voltage | 12,470 V (high-voltage transformer fault forcing fast protective shutdown) |
| UPS bridge performance | Batteries began failing after only ~4 minutes vs their intended ~10-minute bridge |
| Facility generation | 10 generators (inclusive of redundant units), rated to carry PDX-04 at full load — all taken offline by one protective trip |
| Total-power-loss window | 11:44–12:01 UTC (2 Nov) — every DC customer lost power |
| Generator restart delay | Generators restarted 12:48 UTC — ~1h04m after routers dropped / ~1h after UPS began failing |
| Operator notification lag | ~44 minutes (Cloudflare routers offline 11:44 → first Flexential notice 12:28 UTC) |
| Single-site blast radius | >1/3 of the high-availability cluster machines + Cloudflare's largest analytics cluster housed in PDX-04 |
| Hard single-site dependency | Kafka + ClickHouse (log processing) existed ONLY in PDX-04, with dependents in the HA cluster |
| Total incident duration | ~40h41m (11:44 UTC Nov 2 → 04:25 UTC Nov 4); CEO characterized customer pain as 'the last 36 hours' |
| Raw log availability | Down for most customers for the entire duration of the incident (total, not intermittent) |
Magnitude profile
Blast radius (8) and duration (8) dominate: a single-site power loss removed >1/3 of the high-availability cluster plus the largest analytics cluster and broke services meant to survive one-site loss, over a ~36-40 hour window running into a second calendar day. Users score (7) reflects a global control-plane/analytics/raw-log outage that nonetheless spared the data plane (traffic kept flowing). Financial (5) is a placeholder — no dollar impact is disclosed in the primary source; the reputational cost was high (public CEO apology). Scores are structural, not from a verified customer count, which the primary postmortem does not provide.
Sequence of events (SOE)
- TRIGGER Portland General Electric (PGE) has an unplanned maintenance event affecting one of the two independent power feeds into Flexential PDX-04; the building loses one utility feed.
- MITIGATION Flexential powers up on-site generators to supplement the lost feed; facility now runs on generators plus the one remaining utility feed.
- DETECTION Notification/observability gap: Flexential does not tell Cloudflare it failed over to generators, and none of Cloudflare's observability tools detect the power-source change.
- CASCADE A ground fault occurs on a PGE transformer at PDX-04 — believed (unconfirmed by Flexential/PGE) to be the step-down transformer for the second, still-running feed; high-voltage 12,470 V lines involved.
- CASCADE Fast protective shutdown clears the fault but ALSO shuts down all of PDX-04's generators — a protection-coordination failure taking the remaining utility feed and on-site generation together.
- DETECTION Cloudflare is first alerted to trouble only when the two routers connecting the facility to the world go offline — no prior notification from Flexential.
- IMPACT UPS battery bank rated for ~10 minutes begins failing after only ~4 minutes; depletion inferred from Cloudflare's own equipment failing.
- CASCADE Flexential's access-control system, not powered by the battery backups, goes offline — impeding physical entry for manual generator restart.
- IMPACT With generators not restarted, the UPS batteries run out and ALL customers of the data center lose power (total facility de-energisation).
- CASCADE Logical propagation: Kafka and ClickHouse — hosted only in PDX-04 — go down, taking dependent high-availability-cluster services (control plane + analytics/logging) with them.
- MITIGATION Blast radius contained to the control/analytics plane: Cloudflare's edge network and security services keep serving traffic — the data plane is not impacted.
- DETECTION Flexential sends Cloudflare its first notification of the incident — about 44 minutes after Cloudflare's routers had already gone dark.
- MITIGATION Manual generator restart is hampered by three factors reported by employees (unofficial): offline access control, difficulty, and thin staffing — 'security and an unaccompanied technician who had only been on the job for a week.'
- RECOVERY Flexential manually restarts generators, restoring partial power to the facility.
- CASCADE On attempting to re-energise Cloudflare's circuits, the circuit breakers are found faulty; cause undetermined (ground fault, surge, or pre-existing defect).
- RECOVERY Cloudflare works to fail control-plane and analytics services onto surviving infrastructure while breakers are replaced; data plane continues serving.
- RESTORED Flexential replaces the failed circuit breakers, restores both utility feeds, and confirms clean power to PDX-04 — about 11 hours after the ground fault.
- RESTORED Cloudflare completes full restoration of control-plane and analytics services following clean-power restoration.
Root cause
Contributing factors
- Degraded starting posture: PGE's 08:50 UTC unplanned-maintenance loss of one independent feed left PDX-04 running on generators-plus-one-feed for ~2h50m before the ground fault, so redundancy margin was already half-consumed at the moment of the fault (Cloudflare post-mortem).
- Protection mis-coordination (engineering interpretation of Cloudflare's facts): the 12,470 V ground-fault protection 'also shut down all of PDX-04's generators,' converting a single transformer fault into simultaneous loss of the remaining utility feed AND on-site generation (Cloudflare post-mortem).
- UPS battery ride-through collapse: rated 'approximately 10 minutes' but it 'started to fail after only 4 minutes' (inferred from Cloudflare's own equipment failing), closing the window for manual generator restart — consistent with a degraded/under-capacity string, though Cloudflare does not diagnose the cause (Cloudflare post-mortem).
- Faulty circuit breakers on Cloudflare's circuits, discovered only on re-energisation, cause undetermined — 'We don't know if the breakers failed due to the ground fault or some other surge... or if they'd been bad before' — a possible latent, un-inspected defect (Cloudflare post-mortem).
- Communications-procedure lapse: 'counter to best practices, Flexential did not inform Cloudflare that they had failed over to generator power,' and first notification came only at 12:28 UTC, delaying operator awareness (Cloudflare post-mortem).
- Access-control system not on battery backup went offline, physically impeding entry for manual generator restart during the outage (Cloudflare post-mortem).
- Thin, inexperienced overnight staffing (unconfirmed, per employee accounts Cloudflare could not officially verify): 'security and an unaccompanied technician who had only been on the job for a week,' with no experienced ops/electrical expert on site (Cloudflare post-mortem).
- Cloudflare testing gap: 'we had never tested fully taking the entire PDX-04 facility offline,' so hidden single-site Kafka/ClickHouse dependencies inside the high-availability control plane went undetected (Cloudflare post-mortem).
- Observability gap: none of Cloudflare's monitoring tools could detect that PDX-04's power source had changed to generators, so the first alert was hardware death at 11:44 UTC (Cloudflare post-mortem).
- Governance gap: products reached general availability without formally requiring migration to redundant backends, so redundancy protections worked inconsistently by product — Cloudflare: 'That was a mistake' (Cloudflare post-mortem).
Correction of errors (COE)
- Ensure the Cloudflare control plane and analytics can survive the loss of any single data center; remove PDX-04 as a single point of failure for Kafka, ClickHouse, and dependent services.
- Test failover by fully taking core data-center facilities offline (disaster-recovery / chaos drills), the scenario never previously exercised.
- Enforce a governance requirement that no product goes generally available without redundant, multi-region backends.
- Improve observability to detect facility power-source changes and abnormal power states, and require immediate tenant notification on generator failover.
- Facility remediation: replace/inspect faulty breakers, verify UPS battery capacity, put access control on backup power, and review generator protection coordination.
Lessons learnt
- 'Highly available' by architecture is not the same as highly available by test — Cloudflare had never exercised a full-facility loss, so a hidden single-site dependency (Kafka/ClickHouse) survived undetected until a real outage exposed it.
- A single physical fault can defeat logical redundancy: a contained PDX-04 power loss cascaded into the control plane because critical backends lived in only one location.
- Protection that is electrically correct can still be operationally catastrophic — fault-clearing that also shuts down all standby generators turns a single transformer fault into total facility de-energisation.
- Observability must cover the power chain, not just servers: with no visibility into the utility-to-generator failover, the operator's first signal was routers dying, forfeiting any reaction window.
- Vendor communication is a resilience control — a facility operator failing to notify a tenant of generator failover (and a ~44-minute delay before first notice) blinds the tenant precisely when situational awareness matters most.
- Ride-through assumptions must be verified, not trusted: a UPS 'supposedly sufficient for ~10 minutes' delivering ~4 minutes erased the time needed for manual generator restart.
- Recovery paths have their own single points of failure — an access-control system off the battery backup can physically block the very people who must restart generators.
Improvements & remediation
- Eliminate single-site logical dependencies: relocate/replicate Kafka and ClickHouse (and any control-plane backend) across multiple regions so no one facility can take down the high-availability control plane (Cloudflare committed remediation).
- Institute full facility-offline disaster-recovery drills — actually test 'taking the entire PDX-04 facility offline,' the exact scenario Cloudflare admitted it had never tested.
- Enforce a redundancy-governance gate: no product reaches general availability without formally requiring migration to redundant, multi-region backends, closing the inconsistency Cloudflare called 'a mistake.'
- MAINTENANCE — routine UPS battery capacity/load testing and proactive replacement, so a bank rated ~10 minutes cannot silently degrade to ~4 minutes of real ride-through.
- MAINTENANCE — periodic protective-device and circuit-breaker inspection/testing on tenant circuits, to catch latent breaker defects like those 'discovered to be faulty' only on re-energisation.
- MAINTENANCE / COMMISSIONING — review protection coordination and selectivity so a utility-side ground fault cannot trip all on-site generators; the standby generation must ride through exactly this fault class.
- SAFETY / RESILIENCE — power the data-center access-control system from UPS/battery backup so responders can physically reach generators during an outage instead of being locked out.
- SAFETY / STAFFING — guarantee qualified electrical/operations personnel on every overnight shift rather than 'security and an unaccompanied technician... on the job for a week.'
- Add power-source-state observability: monitoring that detects utility-to-generator failover and abnormal power conditions, so operators are not blind until hardware dies.
- Codify a tenant-notification SLA requiring the facility operator to alert customers immediately on generator failover or any power event, fixing the lapse where Flexential 'did not inform Cloudflare.'
Comprehensive analysis
Not a fire — an electrical de-energisation cascade
No public source describes any combustion, fire-detection actuation, suppression discharge, evacuation, or fire-service response. The event is a power-loss cascade: a 12,470 V ground fault on a PGE transformer at Flexential PDX-04, protection that cleared the fault but 'also shut down all of PDX-04's generators,' and a UPS bank that failed in ~4 minutes against a ~10-minute rating. Treating this as a fire would misdiagnose the entire failure spine; the correct forensic frame is protection coordination, ride-through, and manual restart.
Redundancy consumed before the initiating fault
The facility was already degraded when the ground fault struck. PGE's 08:50 UTC unplanned maintenance had removed one of two independent feeds, and Flexential was supplementing with generators. For ~2h50m PDX-04 ran on generators plus a single feed — half its utility redundancy already gone. The ground fault then hit the transformer for that remaining feed, and the protective trip took the generators too. A latent maintenance window on the utility side thus set up the conditions for a single fault to become total loss.
Physical single-site fault, logical single-site dependency
The most consequential lesson is architectural. Physically, only PDX-04 went dark and Cloudflare's data plane kept serving traffic. But the control plane and analytics collapsed because Kafka and ClickHouse existed only in PDX-04 while high-availability-cluster services depended on them. A system marketed as highly available carried an untested single-region dependency. Cloudflare is candid that products reached general availability 'without formally requiring' redundant backends and that 'we had never tested fully taking the entire PDX-04 facility offline' — 'That was a mistake.'
Recovery friction: observability, communication, and staffing
Cloudflare had no signal of the power-source change and was first alerted only when its routers died at 11:44 UTC; Flexential's first notification came at 12:28 UTC. Manual generator restart — the only path, since automatic restart did not occur — was slowed by an access-control system off the battery backup and, per unconfirmed employee accounts, an overnight shift of 'security and an unaccompanied technician who had only been on the job for a week.' Partial power returned at 12:48 UTC but Cloudflare's breakers were found faulty; clean power was not confirmed until 22:48 UTC, roughly 11 hours after the fault.
Evidence quality and disclosed gaps
Every load-bearing claim traces to Cloudflare's own official post-mortem, an authoritative vendor account of its own outage. Its limits are disclosed and preserved here: Cloudflare could not get Flexential/PGE confirmation on the transformer, does not know whether the breakers were pre-faulted, inferred UPS depletion from its own failing gear, and flags the staffing detail as unconfirmed employee testimony. No independent regulatory, court, or PGE/Flexential engineering record is public, so facility-side findings rest on one party's narrative and should not be read as regulator-verified.
Technical deep-dive
References & provenance
- official-postmortem Post Mortem on Cloudflare Control Plane and Analytics Outage“Ground faults with high voltage (12,470 volt) power lines are very bad... Unfortunately, in this case, the protective measure also shut down all of PDX-04's generators.”https://blog.cloudflare.com/post-mortem-on-cloudflare-control-plane-and-analytics-outage/
- news DCD: Cloudflare suffers power outage at Flexential Oregon data center“(Secondary trade-press coverage of the Flexential/Hillsboro power outages; not re-verified from primary source this pass — source returned 403 and a facility-name discrepancy ('PDX01' vs Cloudflare's 'PDX-04') plus later second-outage details remain unverified. Treated as screening-only.)”https://www.datacenterdynamics.com/en/news/cloudflare-suffers-second-power-outage-at-flexential-data-center-in-oregon-in-six-months/
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-01.