← All incidents
Incident dossier · Rank #21

Microsoft Azure / Microsoft 365 Global WAN Outage — IGP-Purge Command Cascade (Jan 25 2023)

Microsoft Azure 2023-01-25 5h 35m core impact Networkconfig-errorchange-management

A single IGP-database-purge command, added to a WAN change procedure that was never re-tested, executed network-wide on one router vendor's platform in Madrid — forcing every IGP-joined router in Microsoft's global backbone (AS 8075) to recompute adjacency and forwarding tables and BGP to re-validate all internet prefixes. The resulting packet loss took Azure and most of Microsoft 365 offline worldwide for roughly 5.5 hours, with a second wave triggered 33 minutes later when the same command was run on a second Madrid router.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2023-01-25)Trigger · Network2023-01-252023-01-25Primary fault at Microsoft Azure — Global WAN backbone (Madrid capacity-expansion node, AS 8075)Microsoft AzureGlobal WAN backbone (Madrid capacity-expansion node, AS 8075)Global WAN backboneDownstream service degraded by the fault: Microsoft TeamsMicrosoft TeamsDownstream service degraded by the fault: Exchange Online / OutlookExchange Online / OutlookDownstream service degraded by the fault: SharePoint OnlineSharePoint OnlineDownstream service degraded by the fault: OneDrive for BusinessOneDrive for Business+8 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Microsoft Azure
Data center
Global WAN backbone (Madrid capacity-expansion node, AS 8075)
Location
Madrid, Spain, n/a — backbone WAN, not a single AZ
Date
2023-01-25

Impact & scale

Users affected
Not precisely disclosed by Microsoft. Any user served by the affected infrastructure worldwide; majority in Asia-Pacific and EMEA as it hit core business hours. Third-party Downdetector logged thousands of reports — more than 3,900 in India and over 900 in Japan (NPR).
Financial
Not disclosed by Microsoft or any regulator.
Scope
Sev-A global multi-service outage
Services / systems down
  • Microsoft Teams
  • Exchange Online / Outlook
  • SharePoint Online
  • OneDrive for Business
  • Microsoft Graph
  • Power BI / Power Platform
  • Microsoft 365 admin center
  • Microsoft Intune
  • Microsoft Defender for Cloud Apps / Identity / Endpoint
  • Azure client connectivity
  • Azure ExpressRoute
  • Azure Government (reported by news coverage; not itemised in the M365 PIR)

Impact data & metrics

Core incident duration5h 35m (07:08–12:43 UTC)
Bulk / peak impact window~90 minutes
Full WAN prefix reconvergence time~1 hour 40 minutes
Detection latencyAlerts to on-call within ~5 min (07:08 → 07:11 UTC)
Delay to second (repeat) wave33 minutes after first command
SharePoint/OneDrive request drop~10% reduction in worldwide RPS, 07:08–09:50 UTC vs prior week
Routers that failed automatic recovery2
Autonomous System affectedAS 8075 — multiple prefixes (/24 and summary /12) withdrawn then readvertised
Residual Outlook-on-web full resolution2023-01-26 03:50 UTC (EX502694)
Third-party outage reports (Downdetector)>3,900 reports in India, >900 in Japan; thousands worldwide

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 6Outage duration (0–10) — how long service was degraded/down. — scored 6/10.Blast 9Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 9/10.
Magnitude 7.5 = blast 9×0.35 + users 8×0.25 + financial 6×0.20 + duration 6×0.20 (sub-scores 0–10 · weighted composite)

Blast radius maximal: the fault was in the global WAN itself (AS 8075), so it degraded client-to-Azure, datacenter-to-datacenter and ExpressRoute paths simultaneously (all three named in Microsoft PIR MO502273), taking down essentially the whole Microsoft 365 suite. Azure Government impact was reported by news coverage but is not itemised in the M365 PIR. Duration 5h35m core (07:08–12:43 UTC) with residual Outlook-on-the-web impact tracked separately (EX502694) into Jan 26 03:50 UTC. Financial score inferred from global productivity impact — no USD figure was disclosed by Microsoft or any regulator, so this is an estimate, not a sourced number.

Sequence of events (SOE)

Phased sequence of events2023-01-25 07:08 UTC · TRIGGER — A network engineer performing a planned task to add WAN capacity in Madrid modifies the IP address on a new router and runs a command to purge the IGP database. On this router vendor's platform the command's default action is global, not local, so it executes across every IGP-joined router.TRIGGER2023-01-25 07:08 U2023-01-25 07:08 UTC · CASCADE — All IGP-joined routers in Microsoft's global WAN are ordered to recompute their IGP topology (adjacency and forwarding) databases; during the recomputation the routers cannot correctly forward packets.CASCADE2023-01-25 07:08 U2023-01-25 ~07:10 UTC · CASCADE — BGP routers (AS 8075) begin readvertising and re-validating internet prefixes as a second-order effect; ThousandEyes observes multiple Microsoft prefixes withdrawn completely then almost immediately readvertised, precipitating global route-table instability and packet loss.CASCADE2023-01-25 ~07:10 U2023-01-25 07:09–07:22 UTC · IMPACT — First wave of SharePoint Online / OneDrive availability loss (of three waves: 07:09–07:22, 07:42–07:47, 08:27–08:29) as User Front End components lose internal connectivity to backend SQL infrastructure.IMPACT2023-01-25 07:09–07:22 U2023-01-25 07:11 UTC · DETECTION — Microsoft monitoring alerts on-call engineers of a networking issue (within ~3–5 minutes of the command); the team initially begins reviewing DNS configurations.DETECTION2023-01-25 07:11 U2023-01-25 07:27 UTC · DETECTION — Microsoft publishes incident MO502273 to the Service Health Dashboard and status.office.com after the admin center is affected.DETECTION2023-01-25 07:27 U2023-01-25 07:38 UTC · MITIGATION — SPO and ODB are failed off Azure Front Door routing to test for relief; this does not correct the problem.MITIGATION2023-01-25 07:38 U2023-01-25 07:40 UTC · DETECTION — Monitoring confirms the fault is inside the Microsoft WAN, not DNS — redirecting the investigation to the backbone.DETECTION2023-01-25 07:40 U2023-01-25 ~07:41 UTC · CASCADE — Because the engineer was never informed of the alerts (the modified SOP omitted the checks), the same command is run on a second Madrid router ~33 minutes after the first, creating a second wave of connectivity failures across the network.CASCADE2023-01-25 ~07:41 U2023-01-25 07:44 UTC · MITIGATION — Reports confirm multiple services impacted; out of caution Azure reverts a recently deployed IPv6 change suspected as the cause — later confirmed unrelated.MITIGATION2023-01-25 07:44 U2023-01-25 08:10–08:20 UTC · RECOVERY — Automated WAN recovery features begin restoring the network just after 08:10; at 08:20 Microsoft identifies the WAN change as the source of the incident.RECOVERY2023-01-25 08:10–08:20 U2023-01-25 09:00 UTC · RECOVERY — Telemetry shows the majority of network devices recovered; Outlook desktop-client connectivity recovers around this time.RECOVERY2023-01-25 09:00 U2023-01-25 09:35 UTC · RECOVERY — WAN self-recovers and equipment stabilizes; full prefix reconvergence took roughly 1 hour 40 minutes. Two routers fail to recover automatically, leaving residual intermittent packet loss.RECOVERY2023-01-25 09:35 U2023-01-25 09:40–09:50 UTC · RECOVERY — Microsoft Teams recovers at 09:40; SharePoint Online and OneDrive for Business recover by 09:50; some customers begin reporting recovery at 09:51.RECOVERY2023-01-25 09:40–09:50 U2023-01-25 10:11 UTC · IMPACT — Engineers identify additional region-specific packet loss in the India region during the tail of recovery.IMPACT2023-01-25 10:11 U2023-01-25 11:04–11:30 UTC · MITIGATION — Microsoft begins reviewing options to restart affected application pools to clear residual downstream M365 impact; by 11:30 telemetry shows impact no longer occurring for most customers (mailbox application-pool restarts were executed later, ~13:30 UTC).MITIGATION2023-01-25 11:04–11:30 U2023-01-25 12:43 UTC · RESTORED — Majority of Microsoft 365 services recovered, overall service stable, and packet loss returned to normal levels; only a small residual for Outlook on the web remains.RESTORED2023-01-25 12:43 U2023-01-25 14:10 UTC · RESTORED — MO502273 closed as resolved after extended monitoring; the residual Outlook-on-the-web impact is tracked separately under EX502694.RESTORED2023-01-25 14:10 U2023-01-26 03:50 UTC · RESTORED — Residual Exchange Online impact (a BitLocker / Shared Cache race condition on CAFE components after machine restarts) is fully resolved under EX502694.RESTORED2023-01-26 03:50 U

Root cause

Immediate mechanism: At 07:08 UTC on 25 January 2023, a network engineer was adding capacity to Microsoft's global WAN in Madrid. As part of integrating the new routers into the IGP (Interior Gateway Protocol, which connects all routers inside Microsoft's WAN) and BGP (Border Gateway Protocol, which distributes internet routing) domains, the modified procedure included a command to purge the IGP database. That command is not vendor-neutral: on routers from two of Microsoft's manufacturers it acts only on the local device, but on the third manufacturer's platform — the one being changed — its default action is global, ordering every IGP-joined router in the entire backbone to recompute its IGP topology database. While the routers recomputed adjacency and forwarding tables they could not forward packets correctly, so this single command silently detonated across the whole global WAN, degrading client-to-Azure connectivity, datacenter-to-datacenter connectivity, and ExpressRoute simultaneously. The IGP storm then forced BGP (AS 8075) to withdraw and re-validate internet prefixes; ThousandEyes observed Microsoft prefixes being withdrawn and near-immediately readvertised, producing global route instability and heavy packet loss. Latent / organisational root: This was fundamentally a change-management and safeguard failure, not merely "human error." Microsoft's standard operating procedure for this class of change is a four-step gate — validation in the Open Network Emulator (ONE), lab testing, a documented Safe-Fly Review with roll-out/roll-back plans, and Safe-Deployment limiting access to one device at a time. The SOP had been edited shortly before the event to fix problems seen in previous runs, but the edited procedure was never re-tested and shipped without proper pre- and post-checks. Two defences that should have caught the command both had gaps: (1) the real-time Authentication, Authorization and Accounting (AAA) system maintains a block-list of globally-impactful commands, but this command's global default behaviour on that specific router model had never been discovered during the high-impact-command evaluation for that platform, so it was never added to the block-list; and (2) the "one device at a time" Safe-Deployment isolation was rendered useless because the single accessible router was still IGP-connected to the whole backbone — device-level isolation gives no protection against a command that propagates through the routing protocol itself. Finally, although monitoring alerted on-call engineers within five minutes, the engineer making the change was never notified (again a consequence of the unqualified SOP edit), so the identical command was executed on a second Madrid router 33 minutes later, causing an avoidable second wave.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

Textbook change-management failure

This is a textbook change-management failure in which the technical trigger (one CLI command) was trivial but the organisational controls that should have contained it all failed in series: an un-re-tested SOP edit, an incomplete AAA block-list, a device-isolation model blind to protocol-level propagation, and an alerting path that excluded the operator.

Maximal WAN blast radius

The blast radius was maximal because the fault lived in the shared global backbone (AS 8075) rather than in any single service or region — it simultaneously degraded internet-to-Azure, DC-to-DC, and ExpressRoute paths, so essentially every network-dependent Microsoft 365 and Azure service saw impact at once.

Two distinct outage waves

the first at 07:08 UTC and an avoidable second ~33 minutes later when the same command ran on the second Madrid router — a direct consequence of the operator not being wired into the alerting the modified SOP dropped.

Recovery gated by scale

automated WAN recovery began just after 08:10 UTC but full prefix reconvergence took ~1h40m, and downstream services could only recover after the WAN stabilised (~09:35 UTC), which is why user-visible M365 recovery stretched to 12:43 UTC.

Long-tail Exchange Online residual

A long-tail Exchange Online residual (Outlook on the web / EAS) was a separate, secondary failure mode — a BitLocker/Shared Cache race on CAFE components triggered by mass machine restarts — tracked under EX502694 and resolved 2023-01-26 03:50 UTC.

Impact figures undisclosed

No USD or precise user-count figure was disclosed by Microsoft or any regulator; third-party Downdetector data (thousands of reports, >3,900 India / >900 Japan) is the only quantified external gauge of user impact.

Technical deep-dive

The trigger was a command to purge the IGP (Interior Gateway Protocol) link-state database, added to the WAN capacity-expansion SOP for integrating new Madrid routers into Microsoft's IGP and BGP domains. On two of Microsoft's three router vendors this purge is scoped to the local device; on the third vendor's platform — the one being provisioned — the command's default action is global, instructing every IGP-adjacent router in AS 8075 to flush and recompute its link-state/topology database. During that network-wide SPF recomputation the routers could not correctly forward transit traffic, producing immediate, backbone-wide packet loss across client-to-Azure, datacenter-to-datacenter, and ExpressRoute paths. The IGP event then cascaded into BGP. Because interior reachability changed underneath it, Microsoft's BGP speakers (AS 8075) began re-advertising and re-validating the internet prefixes they carry. ThousandEyes observed Microsoft prefixes — both specific /24s and summary prefixes up to /12 — being withdrawn completely and then almost immediately re-advertised, which destabilised global routing to Microsoft's ranges and drove significant external packet loss for roughly a 90-minute bulk window. Given the size of the network, restoring reachability to every prefix took about 1 hour 40 minutes. Two safeguards failed to contain it. The real-time AAA system enforces a per-command block-list of globally-impactful operations, but this command's global default on the changed router model had never been surfaced during that platform's high-impact-command evaluation, so it was absent from the list and executed. Separately, Safe-Deployment restricts access to one device at a time; that model assumes damage is bounded by device access, but this command's effect propagated through the IGP adjacency graph rather than through device-to-device access, so single-device access provided zero containment. Compounding it, on-call was alerted within ~5 minutes but the executing engineer was not — a gap introduced by the un-re-tested SOP edit — so the same command was run on the second Madrid router ~33 minutes later, producing a second wave. Automated WAN recovery began just after 08:10 UTC; the WAN self-stabilised by ~09:35 UTC with two routers requiring manual recovery. Downstream, SharePoint/OneDrive front-ends lost connectivity to backend SQL (three sub-waves), Exchange Online connections failed to reach the service, and a separate BitLocker vs Shared Cache race on Exchange CAFE components after machine restarts caused a degraded-cache long tail for Outlook-on-the-web/EAS, resolved under EX502694 at 2023-01-26 03:50 UTC.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-08.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home