Microsoft Azure / Microsoft 365 Global WAN Outage — IGP-Purge Command Cascade (Jan 25 2023)
A single IGP-database-purge command, added to a WAN change procedure that was never re-tested, executed network-wide on one router vendor's platform in Madrid — forcing every IGP-joined router in Microsoft's global backbone (AS 8075) to recompute adjacency and forwarding tables and BGP to re-validate all internet prefixes. The resulting packet loss took Azure and most of Microsoft 365 offline worldwide for roughly 5.5 hours, with a second wave triggered 33 minutes later when the same command was run on a second Madrid router.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Microsoft Azure
- Data center
- Global WAN backbone (Madrid capacity-expansion node, AS 8075)
- Location
- Madrid, Spain, n/a — backbone WAN, not a single AZ
- Date
- 2023-01-25
Impact & scale
- Users affected
- Not precisely disclosed by Microsoft. Any user served by the affected infrastructure worldwide; majority in Asia-Pacific and EMEA as it hit core business hours. Third-party Downdetector logged thousands of reports — more than 3,900 in India and over 900 in Japan (NPR).
- Financial
- Not disclosed by Microsoft or any regulator.
- Scope
- Sev-A global multi-service outage
- Microsoft Teams
- Exchange Online / Outlook
- SharePoint Online
- OneDrive for Business
- Microsoft Graph
- Power BI / Power Platform
- Microsoft 365 admin center
- Microsoft Intune
- Microsoft Defender for Cloud Apps / Identity / Endpoint
- Azure client connectivity
- Azure ExpressRoute
- Azure Government (reported by news coverage; not itemised in the M365 PIR)
Impact data & metrics
| Core incident duration | 5h 35m (07:08–12:43 UTC) |
| Bulk / peak impact window | ~90 minutes |
| Full WAN prefix reconvergence time | ~1 hour 40 minutes |
| Detection latency | Alerts to on-call within ~5 min (07:08 → 07:11 UTC) |
| Delay to second (repeat) wave | 33 minutes after first command |
| SharePoint/OneDrive request drop | ~10% reduction in worldwide RPS, 07:08–09:50 UTC vs prior week |
| Routers that failed automatic recovery | 2 |
| Autonomous System affected | AS 8075 — multiple prefixes (/24 and summary /12) withdrawn then readvertised |
| Residual Outlook-on-web full resolution | 2023-01-26 03:50 UTC (EX502694) |
| Third-party outage reports (Downdetector) | >3,900 reports in India, >900 in Japan; thousands worldwide |
Magnitude profile
Blast radius maximal: the fault was in the global WAN itself (AS 8075), so it degraded client-to-Azure, datacenter-to-datacenter and ExpressRoute paths simultaneously (all three named in Microsoft PIR MO502273), taking down essentially the whole Microsoft 365 suite. Azure Government impact was reported by news coverage but is not itemised in the M365 PIR. Duration 5h35m core (07:08–12:43 UTC) with residual Outlook-on-the-web impact tracked separately (EX502694) into Jan 26 03:50 UTC. Financial score inferred from global productivity impact — no USD figure was disclosed by Microsoft or any regulator, so this is an estimate, not a sourced number.
Sequence of events (SOE)
- TRIGGER A network engineer performing a planned task to add WAN capacity in Madrid modifies the IP address on a new router and runs a command to purge the IGP database. On this router vendor's platform the command's default action is global, not local, so it executes across every IGP-joined router.
- CASCADE All IGP-joined routers in Microsoft's global WAN are ordered to recompute their IGP topology (adjacency and forwarding) databases; during the recomputation the routers cannot correctly forward packets.
- CASCADE BGP routers (AS 8075) begin readvertising and re-validating internet prefixes as a second-order effect; ThousandEyes observes multiple Microsoft prefixes withdrawn completely then almost immediately readvertised, precipitating global route-table instability and packet loss.
- IMPACT First wave of SharePoint Online / OneDrive availability loss (of three waves: 07:09–07:22, 07:42–07:47, 08:27–08:29) as User Front End components lose internal connectivity to backend SQL infrastructure.
- DETECTION Microsoft monitoring alerts on-call engineers of a networking issue (within ~3–5 minutes of the command); the team initially begins reviewing DNS configurations.
- DETECTION Microsoft publishes incident MO502273 to the Service Health Dashboard and status.office.com after the admin center is affected.
- MITIGATION SPO and ODB are failed off Azure Front Door routing to test for relief; this does not correct the problem.
- DETECTION Monitoring confirms the fault is inside the Microsoft WAN, not DNS — redirecting the investigation to the backbone.
- CASCADE Because the engineer was never informed of the alerts (the modified SOP omitted the checks), the same command is run on a second Madrid router ~33 minutes after the first, creating a second wave of connectivity failures across the network.
- MITIGATION Reports confirm multiple services impacted; out of caution Azure reverts a recently deployed IPv6 change suspected as the cause — later confirmed unrelated.
- RECOVERY Automated WAN recovery features begin restoring the network just after 08:10; at 08:20 Microsoft identifies the WAN change as the source of the incident.
- RECOVERY Telemetry shows the majority of network devices recovered; Outlook desktop-client connectivity recovers around this time.
- RECOVERY WAN self-recovers and equipment stabilizes; full prefix reconvergence took roughly 1 hour 40 minutes. Two routers fail to recover automatically, leaving residual intermittent packet loss.
- RECOVERY Microsoft Teams recovers at 09:40; SharePoint Online and OneDrive for Business recover by 09:50; some customers begin reporting recovery at 09:51.
- IMPACT Engineers identify additional region-specific packet loss in the India region during the tail of recovery.
- MITIGATION Microsoft begins reviewing options to restart affected application pools to clear residual downstream M365 impact; by 11:30 telemetry shows impact no longer occurring for most customers (mailbox application-pool restarts were executed later, ~13:30 UTC).
- RESTORED Majority of Microsoft 365 services recovered, overall service stable, and packet loss returned to normal levels; only a small residual for Outlook on the web remains.
- RESTORED MO502273 closed as resolved after extended monitoring; the residual Outlook-on-the-web impact is tracked separately under EX502694.
- RESTORED Residual Exchange Online impact (a BitLocker / Shared Cache race condition on CAFE components after machine restarts) is fully resolved under EX502694.
Root cause
Contributing factors
- An SOP was edited to fix prior-run problems but was never re-validated in ONE/lab and shipped without proper pre- and post-checks, leaving an error in the live procedure.
- The IGP-purge command has different behaviour across the three router vendors — local on two, global on the third — and that global default was never characterised in the high-impact-command review for the changed platform.
- The AAA command block-list did not include this command for this router model, so the real-time guardrail let a globally-impactful command run.
- Device-level Safe-Deployment isolation (one router at a time) provided no protection because the isolated router was still joined to the backbone via IGP — the blast propagated through the protocol, not through device access.
- On-call alerting fired within ~5 minutes but the executing engineer was never informed, so the same destructive command was run on a second router 33 minutes later, creating a second wave.
- Timing amplified impact: the change hit during core business hours across Asia-Pacific and EMEA.
- Two routers could not auto-recover, and downstream M365 services (SPO/ODB, EXO) needed the WAN restored before they could recover, extending user-visible impact well past the WAN's ~09:35 self-recovery.
- Monitoring generated high-volume per-farm/per-feature alerts (SPO/ODB and Teams) rather than quickly detecting a single wider service-wide event, slowing incident triage.
Correction of errors (COE)
- Audit and block similar commands that can have widespread impact across all three router vendors for all WAN router roles.
- Publish real-time visibility of approved-automated, approved-break-glass, and unqualified device activity so on-call engineers can see who is changing what on network devices.
- Implement regular, ongoing mandatory operational training and attestation of following all SOPs.
- Prioritise a Change Advisory Board (CAB) review of all SOPs still pending qualification within 30 days, including engineer feedback on SOP viability and usability.
- Deploy a fix for the BitLocker / Shared Cache race condition on Exchange CAFE components after restart (residual EX502694 impact).
- Change detection and routing logic so components in a known-unhealthy state (e.g. degraded Shared Cache) stop receiving production traffic.
- Fix an additional low-memory bug that caused unexpected machine reboots (the reboots that surfaced the Shared Cache race condition).
- Improve SPO/ODB monitoring logic and alerting thresholds to detect unexpected regional Request-Per-Second changes rather than per-farm noise.
- Reduce Teams alerting noise during wide outages and improve incident-management efficiency; improve Teams anomaly-detection latency.
Lessons learnt
- Device-level change isolation is not blast-radius isolation: a command that propagates via the routing protocol (IGP) defeats 'one device at a time' safety because the isolated device is still logically joined to the fabric.
- A destructive command's behaviour must be characterised per vendor/platform; assuming a command is local-scope because it is local on two of three vendors is a fatal generalisation.
- Any edit to a qualified SOP invalidates its qualification — an un-re-tested SOP change is an untested change and must not fly.
- Real-time command guardrails (AAA block-lists) are only as good as their coverage; the high-impact-command evaluation missed this platform's global default, so the guardrail silently allowed it.
- Operators executing changes must receive the same alerts as on-call, or they will repeat a destructive action (here, a second router 33 minutes later) before anyone reaches them.
- Monitoring that alerts on many narrow symptoms (per-farm/per-feature) instead of the wider event slows detection of a single systemic cause.
- Downstream services that depend on the network cannot recover ahead of it; WAN-level faults set a hard floor on M365 recovery time regardless of app-tier health.
Improvements & remediation
- Extend the AAA high-impact-command block-list to cover globally-scoped commands across all three router vendors and all WAN router roles (committed Feb 2023).
- Give on-call engineers real-time visibility of automated, break-glass, and unqualified device activity so repeat destructive actions are caught before re-execution (committed Feb 2023).
- Mandatory recurring operational training plus attestation of SOP adherence, and a 30-day CAB re-qualification sweep of all SOPs still pending qualification (committed Feb–Apr 2023).
- Treat any SOP edit as de-qualifying: force ONE/lab re-validation and pre/post-checks before the procedure can be scheduled.
- Resiliency fixes in Exchange Online: fix the BitLocker/Shared Cache race on CAFE, stop routing production traffic to known-unhealthy components, and fix the low-memory reboot bug (committed Feb–Mar 2023).
- Tune SPO/ODB and Teams monitoring to detect wide regional/service-wide events quickly and cut alert noise during large outages (committed Mar–May 2023).
Comprehensive analysis
Textbook change-management failure
This is a textbook change-management failure in which the technical trigger (one CLI command) was trivial but the organisational controls that should have contained it all failed in series: an un-re-tested SOP edit, an incomplete AAA block-list, a device-isolation model blind to protocol-level propagation, and an alerting path that excluded the operator.
Maximal WAN blast radius
The blast radius was maximal because the fault lived in the shared global backbone (AS 8075) rather than in any single service or region — it simultaneously degraded internet-to-Azure, DC-to-DC, and ExpressRoute paths, so essentially every network-dependent Microsoft 365 and Azure service saw impact at once.
Two distinct outage waves
the first at 07:08 UTC and an avoidable second ~33 minutes later when the same command ran on the second Madrid router — a direct consequence of the operator not being wired into the alerting the modified SOP dropped.
Recovery gated by scale
automated WAN recovery began just after 08:10 UTC but full prefix reconvergence took ~1h40m, and downstream services could only recover after the WAN stabilised (~09:35 UTC), which is why user-visible M365 recovery stretched to 12:43 UTC.
Long-tail Exchange Online residual
A long-tail Exchange Online residual (Outlook on the web / EAS) was a separate, secondary failure mode — a BitLocker/Shared Cache race on CAFE components triggered by mass machine restarts — tracked under EX502694 and resolved 2023-01-26 03:50 UTC.
Impact figures undisclosed
No USD or precise user-count figure was disclosed by Microsoft or any regulator; third-party Downdetector data (thousands of reports, >3,900 India / >900 Japan) is the only quantified external gauge of user impact.
Technical deep-dive
References & provenance
- official-postmortem Microsoft 365 Customer-Ready Post Incident Report — MO502273 (Report Date Feb 6, 2023)“At 7:08 AM UTC a network engineer was performing an operational task to add network capacity to the global Wide Area Network (WAN) in Madrid... This change added a command to purge the IGP database – however, the command operates differently based on router manufacturer. Routers from two of our manufacturers limit execution to the local router, while those from a third manufacturer execute across all IGP joined routers.”https://www.msxfaq.de/cloud/marketing/mo502273_postincidentreport.pdf
- official Azure Status History — related Azure DNS/WAN PIR, Tracking ID VSG1-B90“For further information, the Azure PIR can be found at status.azure.com with Tracking ID VSG1-B90 (referenced within MO502273).”https://status.azure.com/en-us/status/history/
- analysis Microsoft Outage Analysis: January 25, 2023 — ThousandEyes“ThousandEyes observed a significant number of BGP route changes for prefixes advertised by Microsoft's AS 8075 beginning at just after 07:10 AM UTC... The bulk of the incident lasted approximately 90 minutes.”https://www.thousandeyes.com/blog/microsoft-outage-analysis-january-25-2023
- news-analysis Global Microsoft cloud-service outage traced to rapid BGP router updates — Network World“Multiple Microsoft BGP prefixes were withdrawn completely and then almost immediately readvertised... The bulk of the service disruptions lasted approximately 90 minutes, although ThousandEyes said it spotted residual connectivity issues the following day.”https://www.networkworld.com/article/971873/global-microsoft-cloud-service-outage-traced-to-rapid-bgp-router-updates.html
- news Microsoft applications like Outlook and Teams were down for thousands of users — NPR“Data from Downdetector showed more than 3,900 incidents in India and over 900 in Japan... Microsoft confirmed in a statement to NPR that the outage was a result of a network change and not outside actors.”https://www.npr.org/2023/01/25/1151279866/outlook-teams-sharepoint-outage-microsoft-365
- news Massive Microsoft 365 outage caused by WAN router IP change — BleepingComputer“The outage was caused by a WAN (wide area network) router IP address change that led to packet-forwarding issues between all other WAN routers.”https://www.bleepingcomputer.com/news/microsoft/massive-microsoft-365-outage-caused-by-wan-router-ip-change/
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-08.