← All incidents
Incident dossier · Rank #5

Google Cloud Global Service Control / IAM Outage — June 12, 2025

Google Cloud 2025-06-12 7h 27m core impact Software

A latent, unguarded code path in Google Cloud's Service Control quota system — added on 2025-05-29 without error handling or feature-flag protection — was triggered on 2025-06-12 when a policy change containing unintended blank fields was inserted into the regional Spanner tables Service Control reads for policy data. The blank metadata replicated globally within seconds; regional quota checks pulled the blank fields, exercised the unprotected path, dereferenced a null pointer, and, with no error handling to contain it, the Service Control binaries crashed and entered a global crash loop. Because Service Control gates API authorization, IAM could no longer issue tokens and most Google Cloud APIs failed worldwide for several hours. The failure cascaded into downstream platforms that depend on GCP, most notably Cloudflare, whose Workers KV backend relied on the affected infrastructure.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Software (2025-06-12)Trigger · Software2025-06-122025-06-12Primary fault at Google Cloud — Google Cloud — Service Control quota system (global control plane)Google CloudGoogle Cloud — Service Control quota system (global control plane)Google Cloud — Service ControlDownstream service degraded by the fault: Identity and Access Management (IAM token issuance)Identity and Access ManagementDownstream service degraded by the fault: Cloud BuildCloud BuildDownstream service degraded by the fault: Cloud Key Management ServiceCloud Key Management ServiceDownstream service degraded by the fault: Google Cloud StorageGoogle Cloud Storage+22 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Google Cloud
Data center
Google Cloud — Service Control quota system (global control plane)
Location
USA, global
Date
2025-06-12

Impact & scale

Users affected
Global — customers of dozens of Google Cloud products (~70+, approximate) across every region, plus downstream end-users of dependent platforms (Cloudflare, and widely reported others such as Spotify and Discord). No authoritative aggregate end-user count published; Downdetector peak report counts (e.g. Spotify ~46k) are report tallies, not user counts.
Financial
Not publicly quantified — no official SLA-credit total, insurer, or analyst dollar estimate located.
Scope
SEV-1 / global control-plane outage
Services / systems down
  • Identity and Access Management (IAM token issuance)
  • Cloud Build
  • Cloud Key Management Service
  • Google Cloud Storage
  • Cloud Monitoring
  • Cloud Dataproc
  • Security Command Center
  • Artifact Registry
  • Cloud Workflows
  • Cloud Run
  • BigQuery
  • Cloud Spanner
  • Compute Engine
  • Cloud Firestore
  • Cloud Pub/Sub
  • Cloud SQL
  • Apigee
  • Vertex AI (Gemini API) / Vertex AI Search
  • Looker
  • Memorystore
  • Cloud Functions
  • Cloud Load Balancing
  • App Engine
  • Cloud Console
  • Cloud DNS
  • Google Security Operations

Impact data & metrics

Total incident duration (GCP)~7h27m (10:51–18:18 PDT)
Time to red-button mitigation~40 min from incident begin
Longest region full-resolution (us-central1)up to ~2h40m
Global metadata replication timewithin seconds
Affected Google Cloud products (approximate)~70+ (order-of-magnitude; status page enumerates on the order of 60–70)
Downstream Cloudflare outage window17:52–20:28 UTC; Cloudflare states it 'lasted 2h28m' (window spans ~2h36m by the clock)
Cloudflare Durable Objects (SQLite) / D1 peak error rate22%
Cloudflare AI Gateway peak error rate97%
Cloudflare Realtime TURN / Pages peak error ratenear 100% / ~100%
Cloudflare Stream peak error rate>90%
Cloudflare Workers / Workers for Platforms error rate~2% / ~10%
Cloudflare Workers KV peak error rate~90%
Cloudflare Images batch-upload failure at peak100% (~97% success overall)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 10Users affected (0–10) — breadth of the user/customer population impacted. — scored 10/10.Financial 8Financial impact (0–10) — direct + consequential cost. — scored 8/10.Duration 7Outage duration (0–10) — how long service was degraded/down. — scored 7/10.Blast 10Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 10/10.
Magnitude 9.0 = blast 10×0.35 + users 10×0.25 + financial 8×0.20 + duration 7×0.20 (sub-scores 0–10 · weighted composite)

Global blast radius across all regions and ~70+ products (approximate) with IAM token issuance down (blast/users = 10). Duration ~7h27m end-to-end, meaningful but bounded and recoverable (7). Financial magnitude inferred high from breadth and multi-platform cascade, but no official dollar figure is published (8, low-confidence).

Sequence of events (SOE)

Phased sequence of events2025-05-29 · TRIGGER — Latent ignition: a new quota-policy-check feature is added to Service Control WITHOUT error handling and WITHOUT feature-flag protection. Defect ships to production and lies dormant for ~14 days.TRIGGER2025-05-292025-06-12 ~10:49 AM PDT · TRIGGER — A policy change containing 'unintended blank fields' is inserted into the regional Spanner policy tables Service Control reads; the metadata begins replicating globally within seconds.TRIGGER2025-06-12 ~10:49 AM PD2025-06-12 ~10:51 AM PDT · IMPACT — Service Control hits the null pointer on the blank fields; binaries enter a global crash loop; multiple GCP products (IAM, Cloud Storage and others) begin returning 503 errors and elevated latency.IMPACT2025-06-12 ~10:51 AM PD2025-06-12 ~10:52 AM PDT (17:52 UTC) · CASCADE — Downstream: Cloudflare Workers KV — 'backed by a third-party cloud provider' — loses availability as the Google Cloud storage dependency fails.CASCADE2025-06-12 ~10:52 AM PD2025-06-12 ~10:53 AM PDT · DETECTION — Within ~2 minutes of onset, Google's Site Reliability Engineering team is triaging the incident.DETECTION2025-06-12 ~10:53 AM PD2025-06-12 ~11:01 AM PDT · DETECTION — Within ~10 minutes of onset, the root cause (unguarded null-pointer path on blank policy fields) is identified.DETECTION2025-06-12 ~11:01 AM PD2025-06-12 ~11:16 AM PDT · MITIGATION — The 'red-button' global feature-disable / rollback (suppression analog) is ready ~25 minutes from the start of the crashes.MITIGATION2025-06-12 ~11:16 AM PD2025-06-12 ~11:31 AM PDT · MITIGATION — Red-button rollout completed within ~40 minutes of onset; Service Control begins recovering region by region.MITIGATION2025-06-12 ~11:31 AM PD2025-06-12 ~11:51 AM PDT · DETECTION — First public incident report posted to Cloud Service Health — delayed ~1 hour because the status infrastructure was itself down in the outage (detection channel shared fate with the failure).DETECTION2025-06-12 ~11:51 AM PD2025-06-12 ~12:30 PM PDT (approx.) · RECOVERY — All locations except us-central1 recover following the red-button rollout (clock time derived from onset plus stated recovery offset).RECOVERY2025-06-12 ~12:30 PM PD2025-06-12 early afternoon PDT · MITIGATION — De-energisation/isolation analog: because Service Control lacked randomized exponential backoff, recovering tasks stampede the underlying infrastructure in us-central1 (thundering-herd / 'herd effect'); engineers manually throttle task creation and use multi-regional routing to prevent re-overloading.MITIGATION2025-06-12 early afternoon PD2025-06-12 midday–afternoon PDT · DETECTION — Root cause and mitigation status published/updated by Google on Cloud Service Health once the status infrastructure was restored.DETECTION2025-06-12 midday–afternoon PD2025-06-12 1:28 PM PDT (20:28 UTC) · RECOVERY — Cloudflare Workers KV recovers, ending a 2h28m downstream outage propagated from the Google Cloud storage failure.RECOVERY2025-06-12 1:28 PM PD2025-06-12 (us-central1, up to ~2h 40m after onset) · RESTORED — us-central1 fully resolved after up to ~2h 40m, recovery deliberately throttled to avoid re-overloading infrastructure; primary global incident closes.RESTORED2025-06-12 (us-cPost-incident · RESTORED — Google commits to remediations: fail-open modularization of Service Control, mandatory feature flags disabled-by-default on critical binaries, improved static analysis/error handling, randomized exponential backoff audit, and monitoring that survives when primary products are down.RESTOREDPost-incident

Root cause

The specific "ignition source" was a code change to Service Control — Google Cloud's global, regionally-replicated control plane for API management, authorization and quota enforcement. On May 29, 2025 a new feature was added to Service Control for additional quota policy checks. This is the software analog of an installed-but-defective component: no physical make/model/age/chemistry — the defect was a code path that dereferenced data it never validated. Two latent design flaws shipped with it and lay dormant for 14 days: per the report the change "did not have appropriate error handling nor was it feature flag protected." Because it was not flag-protected and not disabled-by-default, it went live on the critical path with no kill-switch scoped to it; because it lacked error handling it had no graceful-degradation behaviour for malformed input. The exact failure mechanism was a null-pointer dereference triggering a global crash loop. On June 12 a policy change was inserted into the regional Spanner tables that Service Control uses for policies, and that write contained "unintended blank fields." The unguarded new code path read the missing data and, in the report's words, "the null pointer caused the binary to crash." The poison then propagated at machine speed because the data itself was the fault vector: this metadata "was replicated globally within seconds," so the crash occurred globally as each regional deployment ingested the same blank-field policy and crash-looped near-simultaneously, collapsing the API control plane worldwide within seconds to a couple of minutes of the trigger. The latent root is a chain of change-management, validation and architecture lapses rather than any physical-asset or inspection failure. First, a critical binary reached production without a feature flag and without being disabled by default — no scoped, fast rollback existed at ship time. Second, there was no input/schema validation at the policy-write boundary, so a record with blank fields was accepted into the Spanner policy tables that feed a global control plane. Third, globally-replicated policy data had no incremental, canary or staged propagation — it replicated worldwide "within seconds," removing any blast-radius containment. Fourth, the architecture failed CLOSED (crash-loop returning 503s) instead of failing open or degrading gracefully; Google's own remediation concedes this, committing to "modularize Service Control's architecture, so the functionality is isolated and fails open." Fifth, even the recovery path was unsafe: Service Control lacked appropriate randomized exponential backoff, so recovering tasks stampeded the underlying infrastructure in us-central1 (a thundering-herd / "herd effect"), forcing manual throttling and extending that region's recovery to up to ~2h 40m. The provenance of the malformed blank-field policy write — who or what process authored it and why no validation caught it — is not disclosed in the public record.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What happened

On June 12, 2025, at ~10:51 AM PDT, Google Cloud's Service Control — the global, regionally-replicated control plane that handles API management, authorization and quota enforcement — entered a worldwide crash loop and began returning 503 errors across dependent products (IAM, Cloud Storage and dozens of others). The global control-plane impact lasted roughly three hours, with us-central1 taking up to ~2h 40m to fully recover. Because Service Control gates API authorization, the failure presented to customers as a broad, cross-service Google Cloud outage rather than a single-product incident.

Why it happened (root-cause chain)

A new quota-policy-check feature was added to Service Control on May 29, 2025 without error handling and without feature-flag protection, and lay dormant for ~14 days. On June 12 a policy write containing 'unintended blank fields' was inserted into the regional Spanner policy tables Service Control reads. The unguarded code path dereferenced the missing data — 'the null pointer caused the binary to crash.' The malformed metadata 'was replicated globally within seconds,' so every regional deployment ingested it and crash-looped near-simultaneously. The chain is change-management (no flag, not disabled-by-default), validation (no schema check at the write boundary), propagation (no canary/stage), and architecture (fail-closed) failures — not a physical-asset or inspection failure.

Blast radius and cascade

The fault vector was replicated data, making onset uncontainable, and the blast radius crossed the provider boundary. Cloudflare Workers KV — 'backed by a third-party cloud provider' — lost availability for 2h28m (17:52–20:28 UTC), degrading Cloudflare products that depend on KV. Cloudflare characterised the Google failure as the proximate trigger while stating it remains 'ultimately responsible for our chosen dependencies,' underscoring how single-provider dependency concentration exports one operator's control-plane failure into unrelated platforms.

Detection and response

Internal detection was fast: SRE was triaging within ~2 minutes and the root cause was identified within ~10 minutes. A manual 'red-button' global feature-disable was ready ~25 minutes from onset and rollout completed within ~40 minutes, after which all regions except us-central1 recovered. Two response weaknesses stood out: the Cloud Service Health status stack shared fate with the failing system, delaying the first public report by ~1h; and Service Control's lack of randomized exponential backoff caused recovering tasks to stampede us-central1's infrastructure, forcing manual throttling and multi-regional routing and extending that region's recovery to ~2h 40m.

Remediation and accountability

Google committed to modularize Service Control so it fails open, enforce feature flags disabled-by-default on critical binaries, improve static analysis and error handling, audit for randomized exponential backoff, and harden monitoring to survive primary-product outages. Root-cause-adjacent gaps not explicitly ticketed in the public report — schema validation at the policy-write boundary and staged/canaried global metadata propagation — remain the clearest additional safeguards. Cloudflare, on the downstream side, accepted responsibility for its dependency architecture. The provenance of the malformed blank-field write is not disclosed in the public record.

Technical deep-dive

This was a software/control-plane failure, not a physical fire, so the classical forensic axes (detection, suppression, evacuation, emergency services, de-energisation, containment) map to technical analogs, all grounded in the Google Cloud incident report. IGNITION/FUEL: the flammable material was globally-replicated policy metadata; the ignition was a null-pointer dereference in an unguarded, non-flag-protected quota-check code path added May 29. TRIGGER: a June 12 policy write with "unintended blank fields" into the regional Spanner policy tables; because that metadata "was replicated globally within seconds," the poison fanned out to every regional Service Control deployment before any human could react, and each went into a crash loop. DETECTION (fire-detection analog): automated monitoring plus human SRE triage performed well at the technical layer — products began failing at ~10:51 AM PDT, SRE was triaging within 2 minutes, and within ~10 minutes the root cause was identified. But a detection-channel failure existed customer-side: the first public report reached Cloud Service Health only ~1h after the start of the crashes, because the Cloud Service Health infrastructure was itself down in this outage — the smoke detector lost power in the very fire it was meant to report. SUPPRESSION (fire-suppression analog): a manual "red-button" global feature-disable/rollback. It was ready ~25 minutes from the start and rollout was completed within ~40 minutes, after which all locations except us-central1 recovered. Effectiveness was high everywhere except us-central1, where the suppression's own side effect (recovering tasks with no randomized exponential backoff) overloaded the underlying infrastructure — a thundering-herd on restart. DE-ENERGISATION/ISOLATION analog: in us-central1 engineers deliberately throttled task creation (and used multi-regional routing) to minimise impact on the underlying infrastructure, equivalent to manual load-shedding, extending recovery there to up to ~2h 40m. EVACUATION/EMERGENCY SERVICES: not applicable — no personnel or external responders; the operational analog is that Service Control failed CLOSED rather than degrading gracefully, so dependent APIs (IAM, GCS and others) returned 503s globally. CONTAINMENT/SPREAD: uncontainable at onset because the fault vector was replicated data, and the spread crossed the provider boundary — Cloudflare Workers KV, "backed by a third-party cloud provider," lost availability for 2h28m (17:52–20:28 UTC), with Cloudflare stating the proximate cause was a third-party vendor failure while accepting it is "ultimately responsible for our chosen dependencies."

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home