← All incidents
Incident dossier · Rank #24

Alibaba Cloud Global All-Region Outage — Auth Control-Plane Failure

Alibaba Cloud 2023-11-12 3h 27m core impact Softwareconfig-errorcontrol-planeauthentication

On 12 November 2023 Alibaba Cloud's shared global authentication/authorization (Auth) control plane failed, and from 17:44 to 21:11 China Standard Time every global region simultaneously lost console and API access. Auth-dependent services (OSS, OTS, SLS, MNS) went down, crashing consumer apps such as Taobao and DingTalk, while compute/network data planes (ECS, RDS instances) that use their own credentials kept running. Alibaba officially acknowledged only an 'anomaly' in console and API access and published no root-cause post-mortem; the Auth-control-plane mechanism is an independent inference corroborated by the service-failure pattern.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Software (2023-11-12)Trigger · Software2023-11-122023-11-12Primary fault at Alibaba Cloud — Global control plane (Auth/IAM) — all regionsAlibaba CloudGlobal control plane (Auth/IAM) — all regionsGlobal control planeDownstream service degraded by the fault: Object Storage Service (OSS)Object Storage ServiceDownstream service degraded by the fault: Table Store (OTS)Table StoreDownstream service degraded by the fault: Log Service (SLS)Log ServiceDownstream service degraded by the fault: Message Notification Service (MNS)Message Notification Service+2 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Alibaba Cloud
Data center
Global control plane (Auth/IAM) — all regions
Location
Hangzhou, China, Multiple AZs worldwide; East China 1 (Hangzhou) and North China 2 (Beijing) recovered first
Date
2023-11-12

Impact & scale

Users affected
Not quantified by Alibaba; consumer apps Taobao, Xianyu and DingTalk crashed (each with hundreds of millions of users) but no user figure was disclosed
Financial
Not disclosed; only SLA voucher compensation of 25-30% of monthly service fees was offered
Scope
Tier 1 — global multi-region control-plane outage
Services / systems down
  • Object Storage Service (OSS)
  • Table Store (OTS)
  • Log Service (SLS)
  • Message Notification Service (MNS)
  • Cloud consoles / management console
  • Management APIs (OpenAPI)

Impact data & metrics

Outage duration~3.5 hours (207 min), 17:44-21:11 CST
Regions affectedAll global regions simultaneously (East Asia, SE Asia, Middle East, North America)
Fast vs slow recoveryEast China 1 & North China 2 ~1 hour; other AZs ~3 hours
Monthly availability impactDropped to 99.5% for affected products
SLA compensation25-30% of monthly service fees in vouchers
Core services downOSS, OTS, SLS, MNS + consoles & OpenAPI
Follow-on outage (27 Nov 2023)~1 hour 42 min (09:16-10:58 Beijing time), database OpenAPI, 8 regions
Financial impact (USD)Not disclosed
Users affected (count)Not disclosed by Alibaba

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 5Outage duration (0–10) — how long service was degraded/down. — scored 5/10.Blast 9Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 9/10.
Magnitude 7.2 = blast 9×0.35 + users 8×0.25 + financial 5×0.20 + duration 5×0.20 (sub-scores 0–10 · weighted composite)

Blast radius scored 9 because all global regions failed simultaneously (East Asia, Southeast Asia, Middle East, North America) — the signature of a single shared control-plane component, per Vonng. Users score 8: flagship apps Taobao/Xianyu/DingTalk crashed but Alibaba disclosed no user count. Financial 5: only SLA vouchers (25-30% of monthly fees) disclosed, no dollar impact. Duration 5: 3.5 hours (207 min); East China 1 and North China 2 recovered in ~1 hour, other AZs ~3 hours (Vonng).

Sequence of events (SOE)

Phased sequence of events2023-11-12 17:44 CST (UTC+8) · TRIGGER — The shared global Auth/permission control plane begins returning errors; console and OpenAPI calls start failing simultaneously across every Alibaba Cloud region. Vonng (Feng Ruohang) fixes the outage window start at 17:44.TRIGGER2023-11-12 17:44 CS2023-11-12 ~17:44 CST · DETECTION — Alibaba Cloud monitoring detects an 'anomaly' in cloud product console access and API calls across regions — the only cause Alibaba publicly acknowledged.DETECTION2023-11-12 ~17:44 CS2023-11-12 ~17:45 CST · IMPACT — Object Storage Service (OSS) HTTP API calls fail: OSS requires AK/SK/IAM signature verification against the Auth component, which is now unavailable. Exact minute not individually disclosed (interpolated).IMPACT2023-11-12 ~17:45 CS2023-11-12 ~17:50 CST · CASCADE — Table Store (OTS), Log Service (SLS) and Message Notification Service (MNS) fail because they share the same Auth dependency; ECS, RDS running instances and networking keep operating on their own built-in credentials.CASCADE2023-11-12 ~17:50 CS2023-11-12 ~17:50 CST · IMPACT — Cloud consoles and management APIs (OpenAPI) become unusable worldwide; customers cannot log in, view, or manage their resources even where data planes still run.IMPACT2023-11-12 ~17:50 CS2023-11-12 ~18:00 CST · CASCADE — Downstream consumer applications break as they lose object storage and auth: Taobao (shopping), Xianyu/Idle Fish (resale) and DingTalk (work collaboration) briefly crash for users.CASCADE2023-11-12 ~18:00 CS2023-11-12 ~18:00 CST · CASCADE — Dependent payment, messaging and delivery/logistics systems degrade as their storage and notification back-ends stay unreachable.CASCADE2023-11-12 ~18:00 CS2023-11-12 ~18:10 CST · IMPACT — Blast radius confirmed to span all major geographies — East Asia, Southeast Asia, the Middle East and North America — simultaneously, indicating a cross-regional shared component rather than a local fault.IMPACT2023-11-12 ~18:10 CS2023-11-12 ~18:20 CST · MITIGATION — Engineers work to restore the Auth/permission service; Alibaba published no play-by-play, so the specific remediation steps (e.g. rolling back the permission/blacklist change) are not officially disclosed (timestamp interpolated).MITIGATION2023-11-12 ~18:20 CS2023-11-12 ~18:44 CST · RECOVERY — Hangzhou (East China 1, HQ region) and Beijing (North China 2) recover first — in roughly one hour — well ahead of other regions, per Vonng's regional-recovery analysis.RECOVERY2023-11-12 ~18:44 CS2023-11-12 ~19:30 CST · RECOVERY — Remaining availability zones stay impaired; Vonng notes other AZs took about three hours (versus one hour for the two headquarters regions) to come back (timestamp interpolated).RECOVERY2023-11-12 ~19:30 CS2023-11-12 ~20:30 CST · RECOVERY — OSS/OTS/SLS/MNS progressively restored across the remaining global regions as Auth stabilises; exact per-region timestamps not published (interpolated).RECOVERY2023-11-12 ~20:30 CS2023-11-12 21:11 CST · RESTORED — All global regions restored; total outage window 17:44-21:11, about 3.5 hours (207 minutes).RESTORED2023-11-12 21:11 CS2023-11-13 · RESTORED — Alibaba Cloud publicly attributes the incident only to a console/API 'anomaly' and issues no detailed root-cause post-mortem; Vonng states Alibaba 'refuses to publish a post-mortem report.'RESTORED2023-11-132023-11-13 · IMPACT — Aftermath: the 3.5-hour outage pulls affected products' monthly availability down to 99.5%, triggering SLA compensation of 25-30% of monthly service fees in vouchers.IMPACT2023-11-132023-11-27 09:16-10:58 CST · CASCADE — Follow-on outage 15 days later: console and OpenAPI access to database products (PostgreSQL, Redis, MySQL) fails for ~1 hour 42 minutes across eight regions incl. Beijing, Shanghai, Hong Kong and Virginia — a second console/API control-plane failure in a month.CASCADE2023-11-27 09:16-10:58 CS

Root cause

Immediate mechanism (officially acknowledged, thinly): Alibaba Cloud attributed the outage to an "anomaly" in cloud-product console access and API (OpenAPI) calls. It published no detailed root-cause report, so the precise triggering config or command is not officially confirmed. What is verifiable is the blast-radius signature: every global region failed at the same instant, which can only be caused by a single cross-regional shared component. Independent analysis by Feng Ruohang (Vonng), corroborated by which services died and which survived, points to the shared global authentication/authorization (Auth/IAM) control plane as that component. Failure pattern by dependency: services that call Alibaba's HTTP APIs and must present an AK/SK/IAM signature for every request — OSS (object storage), OTS (Table Store), SLS (Log Service), MNS — all failed, because signature verification depends on Auth. Services with their own built-in credentials that do not round-trip through the cloud Auth layer on the data path — ECS instances, RDS instances, networking — kept running. The consoles and management APIs, which authenticate every operator action through the same Auth layer, were globally unusable. This clean split is the strongest evidence that Auth, not power, network or storage hardware, was the failed component. Reconstructed circular-dependency (unconfirmed, flagged as rumor by Vonng): the permission system pushed a blacklist rule; the blacklist was stored on OSS; reading OSS required a permission check; the permission system in turn needed to read OSS — a deadlock in which the control plane could not recover because it depended on the very service it gates. Alibaba never confirmed this, so it must be treated as a plausible reconstruction, not fact. Latent / organisational root: (1) a single globally-shared Auth control plane with no regional isolation or blast-radius containment, so one bad change or deadlock took down every region at once; (2) a circular dependency between the control plane and the object store it protects, with no break-glass path independent of that store; (3) change management around a high-risk shared component the day after the Double 11 peak; and (4) the absence of a published post-mortem, itself an organisational failure that leaves customers unable to assess residual risk. A safeguard that should have caught it: staged/regional rollout of Auth changes with automatic blast-radius limits, plus a bootstrap path for the permission service that does not require the gated storage.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Simultaneous all-region failure

every global region degraded at 17:44 CST, which is only physically possible if a single cross-regional shared dependency failed. That rules out localised power, network or storage-hardware faults and points to a globally-shared control-plane service.

Service-survival split

Auth-gated request-path services (OSS, OTS, SLS, MNS) and all management-plane operations (console, OpenAPI) failed, while data-plane workloads holding their own credentials (running ECS and RDS instances, networking) continued. This is the signature of an authentication/authorization component failure, not a data-plane failure.

Minimal official disclosure

an 'anomaly' in console and API access — and no post-mortem followed, so the exact trigger (a config push, a certificate/permission change, or the rumored blacklist deadlock) remains unconfirmed. The dossier therefore separates verified facts (window, service split, recovery order, compensation) from the inferred mechanism (Auth) and the rumored deadlock (blacklist-on-OSS).

Near-identical follow-on outage

The near-identical follow-on outage on 27 Nov 2023 (database console/OpenAPI across eight regions for ~1h42m) shows the control-plane fragility was systemic rather than a one-off, reinforcing the regional-isolation and change-management findings.

Timing amplified reputational cost

the failure landed the day after the Double 11 peak and took down consumer flagships (Taobao, Xianyu, DingTalk), making an internal cloud fault immediately visible to hundreds of millions of end users.

Technical deep-dive

Alibaba Cloud's request path for HTTP-API services (OSS, OTS, SLS, MNS) requires every call to carry an AK/SK-derived signature that is validated against a central authentication/authorization (Auth/IAM) service before the request is served. When that Auth service began returning errors at 17:44 CST on 12 Nov 2023, signature verification failed everywhere at once, so all Auth-gated services returned errors globally while the console and OpenAPI — which authenticate each operator action through the same layer — became unusable. Crucially, already-running compute and database instances (ECS, RDS) and the network data plane hold their own credentials and do not re-validate against cloud Auth on the data path, which is why they kept serving traffic throughout. Because the Auth control plane was a single globally-shared component with no regional isolation, there was no blast-radius boundary to contain the fault. The unconfirmed but widely-circulated deadlock theory (Vonng flags it as rumor) explains why recovery was slow rather than instantaneous: a newly-pushed permission/blacklist rule was persisted in OSS, but reading that rule required a permission check, and the permission service in turn needed OSS — a circular dependency with no bootstrap path independent of the gated store. The observed recovery order (Hangzhou/East China 1 and Beijing/North China 2 back in ~1 hour, other AZs ~3 hours) is consistent with the two headquarters regions holding a more direct or manually-restorable path to the Auth/OSS bootstrap. The correct architectural fixes are regional/cell partitioning of Auth, a self-contained bootstrap store plus out-of-band break-glass access for the permission service, and staged rollouts with automated blast-radius caps.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-08.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home