← All incidents
Incident dossier · Rank #20

Azure Active Directory Global Authentication Outage (September 28, 2020)

Microsoft Azure 2020-09-28 5h 0m core impact Software

On 28 September 2020 a service update meant only for an internal Azure AD validation test ring crashed on startup and, because of a latent defect in the Safe Deployment Process (SDP) tooling, was pushed straight into global production. The same defect corrupted the deployment metadata so that all deployment rings were targeted at once and the automated rollback could not run, forcing a slower manual recovery. Users were locked out of Azure Portal, Microsoft 365, Teams, Exchange/Outlook and downstream apps for roughly five hours end-to-end. It is a textbook identity-plane blast-radius event: a single canary-scoped change became a worldwide authentication failure.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Software (2020-09-28)Trigger · Software2020-09-282020-09-28Primary fault at Microsoft Azure — Azure Active Directory — global identity/authentication serviceMicrosoft AzureAzure Active Directory — global identity/authentication serviceAzure Active Directory — globalDownstream service degraded by the fault: Azure Active Directory (core)Azure Active DirectoryDownstream service degraded by the fault: Azure PortalAzure PortalDownstream service degraded by the fault: Microsoft 365Microsoft 365Downstream service degraded by the fault: Microsoft TeamsMicrosoft Teams+4 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Microsoft Azure
Data center
Azure Active Directory — global identity/authentication service
Location
USA, global
Date
2020-09-28

Impact & scale

Users affected
Not disclosed by Microsoft — scope described only as 'a subset of customers' across Azure Public and Azure Government clouds; impact quantified via regional authentication success rates rather than an absolute count
Financial
Not disclosed — Microsoft published no SLA-credit total and no third party quantified business loss
Scope
Tier 1 — global identity-plane / single-point-of-failure event
Services / systems down
  • Azure Active Directory (core)
  • Azure Portal
  • Microsoft 365
  • Microsoft Teams
  • Exchange / Outlook email
  • Microsoft Office desktop (some installations)
  • Visual Studio (licensing verification, some installations)
  • OneDrive / SharePoint (related follow-on disruption)

Impact data & metrics

Primary impact duration (to normal operational parameters)~2h58m (21:25 UTC 09-28 → 00:23 UTC 09-29)
End-to-end duration (to residual cleared)~5h (21:25 UTC 09-28 → 02:25 UTC 09-29)
Auth success rate — Europe~81% during outage
Auth success rate — Americas17% at trough, improved to ~37% before mitigation
Auth success rate — Asia72% initially, dropped to ~32% at peak business hours
Auth success rate — Australia~37% during outage
Front-end error symptomHTTP 503 (Service Unavailable) on AAD auth calls
Cloud scopeAzure Public + Azure Government clouds
Absolute users/tenants affectedNot disclosed ('a subset of customers')
Financial impactNot disclosed (no SLA-credit total, no third-party estimate)
Deployment rings targetedAll rings concurrently (intended: 1 validation test ring)
Comms remediation targetInitial customer comms within 15 min of impact (committed)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 6Outage duration (0–10) — how long service was degraded/down. — scored 6/10.Blast 9Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 9/10.
Magnitude 7.5 = blast 9×0.35 + users 8×0.25 + financial 6×0.20 + duration 6×0.20 (sub-scores 0–10 · weighted composite)

Blast radius is the dominant dimension: AAD is the identity control plane for the entire Microsoft cloud, so one backend change removed login for Portal, M365, Teams, Outlook and downstream third-party apps across both Public and Government clouds. Users/financial scores are estimates because Microsoft never published an absolute affected-user count or dollar figure; duration ~5h end-to-end (~2h58m to normal operational parameters).

Sequence of events (SOE)

Phased sequence of events2020-09-28T21:25Z · TRIGGER — Azure AD backend service update intended only for the internal validation test ring is deployed; a latent SDP defect lets it reach production rather than being confined to the validation ring, triggering crash-on-startup of Azure AD backend service instances (software 'ignition source').TRIGGER2020-09-282020-09-28T21:25Z · IMPACT — Global authentication begins failing: customers encounter errors performing authentication for all Microsoft and third-party applications and services dependent on Azure AD, including apps using Azure AD B2C.IMPACT2020-09-282020-09-28T~21:30Z · DETECTION — Automated telemetry (not smoke/heat sensors) detects an unhealthy condition within approximately 5 minutes of impact and engineering is immediately engaged — detection-latency analogue ~5 min.DETECTION2020-09-282020-09-28T21:30–22:00Z · MITIGATION — Concurrent mitigation during troubleshooting: proactive scale-out of some Azure AD services to handle anticipated load once a fix would apply, and failover of certain workloads to a backup Azure AD authentication system.MITIGATION2020-09-282020-09-28T~21:30Z–onward · CASCADE — Designed ring/partition isolation boundary is defeated: because the update was not confined to the customer-data-free validation ring, the crash propagates across geo-distributed active-active partitions globally rather than being contained.CASCADE2020-09-282020-09-28T~21:30Z–02:25Z · IMPACT — Uneven regional blast radius during the incident: Europe 81% auth success, Asia 72% for first 120 min (dropping to a low of 32% at business-hours peak), Australia 37%, Americas 17% (improving to 37% before mitigation).IMPACT2020-09-282020-09-28T~21:30Z–02:25Z · IMPACT — Resilience shielding limits harm: users authenticated before impact were less likely to see errors; Managed Identities for VMs, VM Scale Sets, and AKS held about 99.8% average availability throughout.IMPACT2020-09-282020-09-28T22:02Z · MITIGATION — Root cause established and remediation begun; the automated rollback mechanism ('suppression system') is initiated.MITIGATION2020-09-282020-09-28T22:02Z · CASCADE — Automated rollback FAILS — the same latent SDP defect had corrupted the deployment metadata the rollback depended on, forcing a fallback to manual processes and significantly extending time-to-mitigate (suppression activation-but-failure analogue).CASCADE2020-09-282020-09-28T22:47Z · MITIGATION — Manual mitigation initiated: engineers begin manually updating the service configuration in a path that bypasses the SDP system entirely (de-energisation/isolation analogue — bypassing the failed automated control).MITIGATION2020-09-282020-09-28T23:59Z · RECOVERY — Manual configuration-update operation completes end-to-end across affected backend instances.RECOVERY2020-09-282020-09-29T00:23Z · RECOVERY — Enough backend service instances return to a healthy state to reach normal service operational parameters — effective end of the primary customer-facing impact window.RECOVERY2020-09-292020-09-29T02:25Z · RESTORED — All service instances with residual impact are fully recovered — total restoration.RESTORED2020-09-29Post-incident · RECOVERY — Completed corrective actions: fixed the latent SDP code defect; fixed the rollback system to restore last-known-good metadata against corruption; expanded scope and frequency of rollback operation drills.RECOVERYPost-incidentPost-incident · RESTORED — Remaining committed actions: apply additional SDP backend protections against this issue class; expedite backup authentication system rollout to all key services; onboard Azure AD to a 15-minute automated customer-communications pipeline.RESTOREDPost-incidentN/A · DETECTION — No physical emergency response: no evacuation, no injuries, and no external emergency services (fire/EMS/police) were called — response was entirely Microsoft engineering remediation (fire-brigade/evacuation fields not applicable).DETECTIONN/A

Root cause

TRIGGER. Starting at approximately 21:25 UTC on 28 September 2020, Microsoft deployed a service update to an Azure Active Directory (Azure AD) backend service. The update was intended to reach only an internal validation test ring — an isolated slice containing no customer data, used to catch defects before broad exposure. The update contained a latent code defect that caused the Azure AD backend service to crash on startup. Per Microsoft's RCA (Tracking ID SM79-F88): "A service update targeting an internal validation test ring was deployed, causing a crash upon startup in the Azure AD backend services." Had the update been confined to the test ring, the crash would have been contained to a non-customer-facing slice. ROOT CAUSE — FAILURE OF THE SAFE DEPLOYMENT PROCESS. The safeguard that should have limited blast radius is Microsoft's Safe Deployment Process (SDP), a ringed-rollout system that promotes changes progressively so a defective build is caught in a small ring before reaching production. The true root cause was a latent code defect in the SDP system itself, which delivered the crashing update straight to global production instead of confining it to the validation ring. In Microsoft's words: "A latent code defect in the Azure AD backend service Safe Deployment Process (SDP) system caused this to deploy directly into our production environment, bypassing our normal validation process." The crashing service update was the payload; the SDP defect was the safeguard failure that let it out globally. NOTE ON MECHANISM DETAIL. Some reconstructions state that the SDP mis-targeted rings because it could not "interpret deployment metadata" and therefore "targeted all rings concurrently." Microsoft did not publish that level of internal detail in the canonical RCA, so it is treated here as unverified elaboration rather than an established Microsoft finding. What Microsoft stated is only that a latent SDP code defect deployed the update directly into production, bypassing validation. PROPAGATION / BLAST RADIUS. Azure AD is the shared identity and authentication plane for the Microsoft cloud, so a global crash of its backend translated directly into authentication failures across the ecosystem. Between roughly 21:25 UTC on 28 Sep and 00:23 UTC on 29 Sep 2020, customers may have encountered errors performing authentication operations for all Microsoft and third-party applications and services that depend on Azure AD. Impact was worldwide and varied by region and time of day, but the specific per-region success-rate percentages circulated in some write-ups could not be confirmed against Microsoft's primary RCA and are not asserted here as fact. WHY RECOVERY WAS SLOW — IMPAIRED AUTOMATED ROLLBACK. The standard mitigation for a bad rollout is automated rollback to the last known-good state. That path was impaired for the same underlying reason: the deployment metadata driving the change had been corrupted, so the automated rollback could not cleanly restore a known-good state, and engineers had to fall back to a manual mitigation that bypassed the SDP system. This coupling — the deployment path and its automated recovery path both depending on the corrupted metadata — is why the fastest remedy was unavailable and recovery was extended. Additionally, a separate Azure AD backup authentication system that could have masked authentication failures had not yet been rolled out to all key services (its expedited rollout appears as a committed corrective action). DETECTION, MITIGATION, AND CORRECTIVE ACTIONS. Monitoring detected the unhealthy condition within about 5 minutes and engineering was engaged immediately. Because automated rollback was impaired, engineers escalated to a manual service-configuration update that bypassed the SDP system; recovery was further gated on restarting crashed backend instances and rebuilding a healthy quorum. Enough backend capacity was healthy to meet normal operational parameters by 00:23 UTC on 29 Sep, with residual impact clearing thereafter. Microsoft's committed corrective actions included fixing the SDP code defect, applying additional protections to the backend SDP system, expediting rollout of the Azure AD backup authentication system to all key services, and expanding the scope and frequency of rollback operation drills. (Specific intermediate timeline minute-markers and per-region percentages in the reconstruction were not verifiable against the primary source.)

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What actually failed (and why it was global)

There was no physical event — no fire, no equipment burn, no facility damage. The failed 'component' was the Azure AD Safe Deployment Process (SDP) automation. A routine backend update meant only for an internal validation test ring (which holds no customer data) reached production because of a latent SDP code defect, and the update crashed Azure AD backend service instances on startup. Because the deployment was not confined to the validation ring, the crash propagated across Azure AD's geo-distributed, active-active partitions worldwide, taking down authentication for effectively every Microsoft and third-party service that depends on Azure AD, including Azure AD B2C.

Why a minutes-long fault became a ~3-hour outage

Detection was fast — automated telemetry flagged degradation within about 5 minutes and engineers engaged. The problem was recovery. When automated rollback was initiated at 22:02 UTC, it failed: the same latent SDP defect had corrupted the deployment metadata the rollback depended on. The containment control (phased ring isolation) and the suppression control (automated rollback) shared a single point of failure in the SDP metadata layer, so one defect disabled both. Engineers then had to fall back to a manual configuration update that bypassed SDP entirely (started 22:47 UTC, completed 23:59 UTC), with normal operational parameters restored by 00:23 UTC and residual instances fully recovered by 02:25 UTC.

Blast radius and what limited it

Impact was uneven. Regional authentication success held highest in Europe (81%) and was worst in the Americas (17%, improving to 37% just before mitigation); Asia ran at 72% for the first two hours then fell to a low of 32% as business-hours peak load arrived, and Australia sat at 37%. Two resilience factors blunted harm: users who had authenticated before impact were less likely to see errors, and Managed Identities for VMs, VM Scale Sets, and AKS were shielded by resilience measures at roughly 99.8% average availability. A backup Azure AD authentication system existed and was used to fail over certain workloads, but its coverage was incomplete.

Corrective actions and residual gaps

Microsoft's completed actions removed the ignition source (fixed the latent SDP defect), hardened the suppression control (rollback can now restore last-known-good metadata against corruption), and widened rollback drills. Committed actions add further SDP backend protections, expedite backup-auth rollout to all key services, and automate a 15-minute customer first-notice. One disclosed gap remains open in the record: the PIR does not explain why the latent SDP metadata defect went undetected before deployment, so the adequacy of pre-production test coverage of the SDP automation itself cannot be judged from the published account.

Technical deep-dive

Azure AD is architected as a geo-distributed, active-active service with multiple partitions across multiple data centers around the world, built with isolation boundaries. Changes are meant to flow through a Safe Deployment Process: first a validation ring containing no customer data, then a Microsoft-only inner ring, then production — deployed in phases across five rings over several days. This is the software equivalent of a fire-compartment strategy: blast every ring at once and no firewall is left. The incident is a textbook case of the containment control (ring isolation) and the suppression control (automated rollback) sharing a single point of failure — the SDP metadata layer — so that one latent defect defeated both simultaneously. The failure chain: (1) the SDP metadata handling mis-targeted the deployment, so a validation-ring-scoped update was applied beyond that ring, reaching production; (2) the update contained a fault that crashed Azure AD backend service instances on startup, so global token issuance/authentication began failing at 21:25 UTC; (3) automated monitoring flagged an unhealthy condition within roughly 5 minutes and engineering engaged; (4) parallel mitigations — proactive scale-out of Azure AD services and failover of certain workloads to a backup Azure AD authentication system — were attempted during troubleshooting; (5) at 22:02 UTC root cause was established and automated rollback initiated, but it failed due to the corruption of the SDP metadata; (6) at 22:47 UTC engineers initiated a manual service-configuration update that bypassed SDP, completing by 23:59 UTC; (7) enough backend instances reached normal operational parameters by 00:23 UTC, with residual instances fully recovered by 02:25 UTC. Blast radius was uneven by region and by session state. Regional success rates: Europe 81%, Americas 17% (improving to 37% just before mitigation), Asia 72% for the first 120 minutes then dropping to a low of 32% as peak business-hours traffic hit, Australia 37%. Two resilience factors limited harm: users who had authenticated prior to impact were less likely to experience issues, and Managed Identities for Virtual Machines, VM Scale Sets, and Azure Kubernetes Service were shielded by resilience measures at an average availability of about 99.8% throughout. Scope reached every Microsoft and third-party application dependent on Azure AD, plus applications using Azure AD B2C. No physical forensics apply — no ignition temperature, no suppression agent, no evacuation, no emergency-services call — because there was no facility event; the analogues here are the deployment trigger, telemetry detection, automated rollback (suppression), manual config bypass (de-energisation), and ring isolation (containment).

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home