← All incidents
Incident dossier · Rank #11

Azure Front Door Global Outage: Inadvertent Cross-Build Config Change Trips Latent Data-Plane Crash Bug (29 Oct 2025)

Microsoft Azure 2025-10-29 8h 24m core impact SoftwareHuman error

On 29 October 2025 an inadvertent customer configuration change to Azure Front Door (AFD) — Microsoft's global edge/CDN and traffic-routing layer — generated incompatible configuration metadata because the change sequence spanned two different control-plane build versions. That metadata exposed a latent bug in the AFD data plane, causing nodes to crash during asynchronous processing as the bad config propagated. A deployment protection safeguard activated and blocked further new changes, but it could not stop the already-propagated config: the crash surfaced asynchronously, after propagation, rather than synchronously at deploy time. Within roughly four minutes the config reached a majority of edge sites and within roughly ten minutes the crash had replicated across all edge sites globally. Recovery was complicated because the Last Known Good (LKG) snapshot had itself been updated with the corrupted config; Microsoft recovered by manually editing the latest LKG to strip the offending customer entries — deliberately declining a clean rollback — then redeploying, blocking further customer config propagation, and failing the Azure Portal away from AFD. Customer impact ran from 15:41 UTC on 29 Oct to mitigation at 00:05 UTC on 30 Oct, about 8.5 hours.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Software (2025-10-29)Trigger · Software2025-10-292025-10-29Primary fault at Microsoft Azure — Azure Front Door global edge fleetAzure Front DoorAzure Front Door global edge fleetAzure Front Door global edge fleetDownstream service degraded by the fault: Azure PortalAzure PortalDownstream service degraded by the fault: Microsoft 365Microsoft 365Downstream service degraded by the fault: OutlookOutlookDownstream service degraded by the fault: Xbox LiveXbox Live+4 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Microsoft Azure
Data center
Azure Front Door global edge fleet
Location
Global
Date
2025-10-29

Impact & scale

Users affected
Not quantified by Microsoft. Global blast radius: the crash replicated across all Azure Front Door edge sites worldwide, degrading Microsoft-first-party and customer properties fronted by AFD. Microsoft's authoritative PIR publishes no affected-user or affected-tenant count.
Financial
Not published by Microsoft; no financial figure appears in the authoritative PIR.
Scope
Critical - global edge/control-plane outage affecting first-party and customer traffic
Services / systems down
  • Azure Portal
  • Microsoft 365
  • Outlook
  • Xbox Live
  • Microsoft Copilot
  • Minecraft
  • Azure Communication Services
  • many downstream customer sites fronted by AFD

Impact data & metrics

Total customer-impact window~8 hours 24 minutes (15:41 UTC 29 Oct → 00:05 UTC 30 Oct)
Global propagation / containment failureReached all edge sites ~4 min after customer impact (15:41 → 15:45 UTC); majority of sites by 15:39 UTC
Asynchronous crash-surfacing delay (why the safeguard failed)~5 minutes (config applied 15:36 → data-plane crashes 15:41 UTC)
Detection gap (impact → investigation start)~7 minutes (15:41 → 15:48 UTC), via secondary impact alerts only
Time to first public communication~37 minutes after customer impact (15:41 → status page 16:18 UTC)
Data-plane recovery time at incident~4.5 hours (since improved to ~1 hour; target ~10 minutes by March 2026)
Config-block to full mitigation span~6 hours 35 minutes (config block 17:30 UTC → full mitigation 00:05 UTC)
Config propagation-time reduction (remediation)45 minutes → ~15 minutes (Jan 2026)
Named affected services13 Azure services (incl. Portal, SQL Database, App Service, AD B2C, Databricks, Maps, Media Services, Static Web Apps, Communication Services) plus Microsoft 365, Dynamics 365, Entra ID, Copilot for Security — 'included, but were not limited to'
Control-plane build versions involved in trigger2 concurrent (mismatched) build versions

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 9Users affected (0–10) — breadth of the user/customer population impacted. — scored 9/10.Financial 7Financial impact (0–10) — direct + consequential cost. — scored 7/10.Duration 6Outage duration (0–10) — how long service was degraded/down. — scored 6/10.Blast 10Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 10/10.
Magnitude 8.3 = blast 10×0.35 + users 9×0.25 + financial 7×0.20 + duration 6×0.20 (sub-scores 0–10 · weighted composite)

Blast radius is maximal (10): the crash reached every AFD edge site globally within ~10 minutes, taking down a shared control/edge layer beneath both Microsoft-first-party services and countless customer sites. Users score is very high (9) given global consumer + enterprise reach, though Microsoft published no user count. Duration (6) reflects ~8.5 hours to mitigation — long for a control-plane fault but bounded. Financial (7) is an analyst estimate only: no monetary figure appears in the PIR, so this reflects the evident scale of a global multi-service outage, not a sourced number.

Sequence of events (SOE)

Phased sequence of events15:35 UTC, 29 Oct 2025 · TRIGGER — Incompatible customer-configuration metadata is generated when a valid, non-malicious sequence of changes is processed across two different control-plane build versions (ignition source introduced).TRIGGER15:35 U15:36 UTC · TRIGGER — Corrupt metadata is applied to the pre-production stage of the configuration-propagation pipeline; it passes structural validation because the data-plane crash is asynchronous and has not yet surfaced.TRIGGER15:36 U15:39 UTC · CASCADE — Corrupt metadata reaches the majority of AFD edge sites; the latent data-plane bug is now armed fleet-wide but has not yet crashed instances.CASCADE15:39 U15:41 UTC · IMPACT — Data-plane crashes begin; customers on AFD and Azure CDN experience connection-timeout errors and DNS resolution failures. Official customer-impact-window start.IMPACT15:41 U15:43 UTC · DETECTION — Automatic configuration-protection system activates but is defeated: because the crash surfaced asynchronously, propagation monitoring continued to receive healthy signals and the bad config advanced past the safeguards.DETECTION15:43 U15:45 UTC · CASCADE — Corrupt metadata has replicated across all edge sites globally — global containment failure; the fault reached the full fleet ~4 minutes after customer-impact onset.CASCADE15:45 U15:48 UTC · DETECTION — Engineering incident investigation commences following secondary impact alerts (~7 min after customer impact) — NOT via the config-safety detector, which reported healthy.DETECTION15:48 U16:15 UTC · DETECTION — Examination of AFD configuration changes begins as the investigation narrows toward the config-propagation pipeline as the source.DETECTION16:15 U16:18 UTC · MITIGATION — First public communication posted to the Azure status page (~37 min after customer impact).MITIGATION16:18 U16:20 UTC · MITIGATION — Targeted communications to impacted customers sent via Azure Service Health.MITIGATION16:20 U17:10 UTC · MITIGATION — Engineers begin manually updating the 'Last Known Good' (LKG) configuration to remove the problematic configurations — manual suppression takes over after the automatic safeguard failed.MITIGATION17:10 U17:26 UTC · MITIGATION — Azure Portal 'failed away from Azure Front Door' onto alternate infrastructure so operators and customers retain a management plane during recovery.MITIGATION17:26 U17:30 UTC · MITIGATION — All AFD customer configuration changes/propagation blocked to stop corrupt metadata re-igniting the fleet.MITIGATION17:30 U17:40 UTC · MITIGATION — Deployment of the corrected LKG configuration is initiated across the edge fleet.MITIGATION17:40 U17:50 UTC · MITIGATION — Corrected LKG configuration is available to all edge sites, which begin gradually reloading it; phased recovery deliberately paced to avoid overloading the recovering fleet.MITIGATION17:50 U18:30 UTC · RECOVERY — AFD DNS servers recover — first measurable customer improvement.RECOVERY18:30 U20:20 UTC · RECOVERY — Automatic traffic management is restored across AFD, resuming normal routing behavior.RECOVERY20:20 U00:05 UTC, 30 Oct 2025 · RESTORED — Full mitigation confirmed; official impact-window close (~8h24m total; ~6h35m from the config-change block).RESTORED00:05 U

Root cause

IGNITION SOURCE (root trigger): The spark was not a physical failure but incompatible customer-configuration metadata. Per the Azure PIR (YKYN-BWZ): "A specific sequence of customer configuration changes, performed across two different control plane build versions, resulted in incompatible customer configuration metadata being generated." The changes themselves were valid and non-malicious; the trigger was the INTERACTION of legitimate config changes applied across two coexisting, mismatched Azure Front Door (AFD) control-plane build versions, emitting metadata that neither build version alone would have produced. FUEL / LATENT HAZARD: a pre-existing latent defect in the AFD DATA PLANE. The incompatible metadata, when deployed to edge-site servers, "exposed a latent bug in the data plane." The data-plane code carried a dormant fault that only ignited when fed the cross-build metadata; the control-plane version skew was the accelerant that surfaced it. The exact code-level nature of the latent bug and the specific build-version numbers are NOT disclosed in the record. FAILURE MECHANISM: the incompatibility triggered a crash during ASYNCHRONOUS processing within the data-plane service. The decisive property was asynchrony: the crash surfaced only after the config had already been accepted and propagated (config applied to pre-production 15:36 UTC, data-plane crashes began 15:41 UTC — an ~5-minute latency), so the malformed configuration first passed validation and early-stage health checks looking healthy, then crashed data-plane instances fleet-wide after global propagation. Because AFD deploys customer configuration globally across all edge sites for a consistent experience, there was effectively no containment at ignition — metadata reached the majority of edge sites by 15:39 UTC and replicated across all edge sites globally by 15:45 UTC. Manifest symptoms were connection-timeout errors and DNS resolution failures on AFD and Azure CDN. LATENT ROOT — TESTING/VALIDATION LAPSE: the record names a specific pre-production validation gap: the defect "escaped detection due to a gap in our pre-production validation, since not all features are validated across different control plane build versions." The validation framework never exercised backward/cross-build compatibility of configuration metadata — precisely the fault class that ignited the incident. Compounding this were three design lapses: (1) the safe-deployment/canary safeguard judged health synchronously while the crash was asynchronous, so the protection system activated (15:43 UTC) but was defeated because the crash had not yet registered; (2) configuration processing was not isolated from active traffic-serving instances, so a bad config crashed the serving fleet directly rather than a sacrificial worker; and (3) global-consistency deployment gave the bad config a fleet-wide blast radius with no cell/segment boundary. Data-plane recovery time at the time of the incident was approximately 4.5 hours, itself a recovery-design lapse (since improved to ~1 hour).

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What actually failed: a control-plane/data-plane seam, not a single component

No hardware failed and no malicious or malformed customer request was involved. Individually valid customer configuration changes were processed across two coexisting control-plane build versions during a normal rolling deployment, producing compiled metadata that neither build alone would emit. That metadata was structurally valid enough to enter the propagation pipeline, but it exposed a latent bug in the AFD data plane. The incident is best read as an interaction failure at the seam between two subsystems, armed by a dormant defect that only the cross-build metadata could trigger.

Why the safety system was present yet ineffective

AFD's safe-deployment model assumed a bad configuration would fail fast and synchronously at pre-production/canary, where propagation monitoring would see unhealthy signals and stop advancement. The data-plane crash was asynchronous, surfacing roughly five minutes after the config was accepted (applied 15:36 UTC, crashes 15:41 UTC). The automatic configuration-protection system did activate at 15:43 UTC, but monitoring had continued to receive healthy signals long enough for the config to clear the gates and propagate globally by 15:45 UTC. The safeguard was chasing a hazard that had not yet materialized.

Blast radius and containment

Because AFD intentionally deploys one configuration identically to every edge site to give customers a consistent experience, there was no cell or segment boundary to stop propagation. The corrupt metadata reached the majority of edge sites within ~4 minutes of generation and the entire global fleet by 15:45 UTC. Impact rippled into at least 13 named Azure services (Portal, SQL Database, App Service, AD B2C, Databricks, Maps, Media Services, Static Web Apps, Communication Services and more) plus Microsoft 365, Dynamics 365, Entra ID and Copilot for Security — the PIR notes the list was not exhaustive.

Recovery path and its constraints

Mitigation required stopping re-ignition before repair: engineers began manually rebuilding a Last Known Good configuration at 17:10 UTC, failed the Azure Portal away from AFD at 17:26 so operators kept a management plane, then blocked all AFD config changes at 17:30 to prevent corrupt metadata re-propagating. The corrected LKG was deployed from 17:40 and available to all edges by 17:50, after which the fleet gradually reloaded it. Restoration was deliberately phased to avoid overwhelming a recovering fleet; DNS recovered at 18:30, automatic traffic management at 20:20, and full mitigation was declared at 00:05 UTC on 30 October — an ~8h24m customer-impact window dominated by a ~4.5-hour data-plane recovery time.

Remediation credibility

The corrective set maps cleanly onto each causal layer: synchronous processing and removal of asynchrony attack the crash-latency blind spot; extra rollout stages and longer bake times give delayed faults time to register; worker-process isolation and micro-cell segmentation bound the blast radius; the enhanced cross-build compatibility testing closes the specific validation gap named as the escape reason; and faster recovery/propagation shrink the outage envelope. The dated commitments (Jan–June 2026, with a March 2026 ~10-minute recovery target) make the plan auditable rather than aspirational.

Technical deep-dive

Azure Front Door is a global anycast Layer-7 edge/CDN platform: customer configuration authored in the control plane is compiled to metadata and propagated to data-plane servers on every edge site worldwide, which then serve TLS termination, routing, and DNS resolution. The failure lived at the seam between control plane and data plane. Two control-plane build versions were live concurrently (a normal state during a safe rolling deployment). A customer's sequence of individually valid changes was processed partly by one build and partly by the other, and the resulting compiled metadata was internally inconsistent in a way that neither build alone would emit. This metadata was syntactically valid enough to pass structural validation and enter the propagation pipeline. The pipeline's safety model assumed a bad configuration would fail FAST and SYNCHRONOUSLY at an early stage — pre-production, then canary — where propagation monitoring would see unhealthy signals and halt advancement. The latent data-plane bug violated that assumption: parsing/applying the incompatible metadata triggered a crash inside an ASYNCHRONOUS processing path that faulted only after the config was accepted (config applied to pre-production 15:36 UTC; crashes began 15:41 UTC). In that window the config passed the protection safeguards and propagated while configuration-propagation monitoring continued to receive healthy signals. The safeguard that activated at 15:43 UTC was therefore chasing a hazard that had not yet materialized — sensing the event too late to stop spread. By the time crashes cascaded, the metadata had reached all global edge sites (15:45 UTC), and the data plane was crash-looping across the fleet, producing connection timeouts and DNS resolution failures. Recovery was constrained by architecture. The fix was to restore a Last Known Good (LKG) configuration, but engineers first had to stop the source of re-ignition — blocking AFD customer configuration changes (17:30 UTC) so corrupt metadata could not re-propagate — then push the corrected LKG (manual LKG update begun 17:10; deployment initiated 17:40; available to all edges 17:50, which then gradually reloaded LKG). Restoration was deliberately phased to avoid overwhelming a recovering fleet, and the data-plane recovery path took ~4.5 hours end-to-end at incident time. A critical first-party control surface was failed off the burning fleet (Azure Portal "failed away from Azure Front Door" at 17:26 UTC) so operators retained a management plane during recovery. The remediation set attacks each layer: eliminate asynchrony (removing asynchronous processing from the data plane; enforcing complete synchronous processing of each customer configuration before advancing to production stages); add rollout stages and extend bake time so a delayed crash has time to register before promotion; isolate the blast radius (decoupling configuration processing in data-plane servers from active traffic-serving instances to isolated worker process instances, and introducing micro-cell segmentation); close the test gap (ensuring backwards compatibility with configurations generated across previous build versions); and shrink recovery (data-plane recovery ~4.5h to ~1h, target ~10 min by March 2026; propagation 45 min to ~15 min).

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home