← All incidents
Incident dossier · Rank #6

AWS US-EAST-1 DynamoDB DNS Race Condition Cascades Across the Internet

Amazon Web Services 2025-10-20 14h 32m core impact SoftwareNetwork

A latent race condition in DynamoDB's automated DNS management system left two DNS Enactor processes racing in the us-east-1 region. An unusually delayed Enactor applied an old plan that overwrote a newer one, and a second Enactor's cleanup automation then deleted the plan as stale — instantly removing every IP address for the DynamoDB regional endpoint and leaving clients unable to resolve the service. Because core AWS control planes depend on DynamoDB, the fault cascaded: EC2's DropletWorkflow Manager fell into congestive collapse (insufficient-capacity launch failures), and NLB health-checks began flapping as new instances were brought into service before their network state propagated. DynamoDB DNS was manually corrected by 2:25 AM PDT, but full recovery of dependent services did not complete until 2:20 PM PDT on October 20, 2025 — roughly 14.5 hours after onset. Downdetector logged millions of reports across more than a thousand consumer and enterprise services.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Software (2025-10-20)Trigger · Software2025-10-202025-10-20Primary fault at Amazon Web Services — US-EAST-1 (Northern Virginia)AWSUS-EAST-1 (Northern Virginia)US-EAST-1Downstream service degraded by the fault: Amazon DynamoDBAmazon DynamoDBDownstream service degraded by the fault: Amazon EC2Amazon EC2Downstream service degraded by the fault: Amazon ECSAmazon ECSDownstream service degraded by the fault: Amazon EKSAmazon EKS+6 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Amazon Web Services
Data center
US-EAST-1 (Northern Virginia)
Location
Ashburn, United States, us-east-1
Date
2025-10-20

Impact & scale

Users affected
Millions of end users across 1,000+ downstream internet services; press estimates of affected businesses range ~1,000-2,500 companies (secondary, unreconciled). AWS disclosed no official customer count.
Financial
No official AWS figure. Analyst/press estimates: >US$1B to the global economy, with one industry CEO suggesting the true toll could reach hundreds of billions in lost productivity (Catchpoint, via CNN) — estimates, not measured.
Scope
Sev-1 / Large-scale multi-service regional disruption
Services / systems down
  • Amazon DynamoDB
  • Amazon EC2
  • Amazon ECS
  • Amazon EKS
  • AWS Fargate
  • Amazon Connect
  • AWS STS
  • AWS IAM
  • Amazon Redshift
  • Network Load Balancer (NLB)

Impact data & metrics

Onset2025-10-19 23:48 PDT (06:48 UTC Oct 20)
DynamoDB DNS restored2025-10-20 02:25 PDT
Event end / full recovery2025-10-20 14:20 PDT
Total event duration~14.5 hours (872 min)
DynamoDB API error window (derived)~2h52m (11:48 PM-2:40 AM PDT)
EC2 launch-failure window (derived)~14h (11:48 PM-1:50 PM PDT)
NLB connection-error window (derived)~8h39m (5:30 AM-2:09 PM PDT)
Downstream services affected1,000+ services (Downdetector-tracked)
Internal AWS services affected (press)64 (secondary paraphrase; not an AWS-stated count)
Snapchat Downdetector peak>22,000 reports
Aggregate Downdetector volume (secondary, unreconciled)~4M within 2h; ~6.5M worldwide
Affected businesses (secondary, unreconciled)~1,000-2,500 companies (vs CyberCube ~70,000 orgs)
Estimated economic impact (analyst/press)>US$1B global; up to hundreds of billions (Catchpoint est.)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 9Users affected (0–10) — breadth of the user/customer population impacted. — scored 9/10.Financial 8Financial impact (0–10) — direct + consequential cost. — scored 8/10.Duration 7Outage duration (0–10) — how long service was degraded/down. — scored 7/10.Blast 10Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 10/10.
Magnitude 8.8 = blast 10×0.35 + users 9×0.25 + financial 8×0.20 + duration 7×0.20 (sub-scores 0–10 · weighted composite)

Blast radius maximal: a single region's DNS fault for one foundational service (DynamoDB) propagated through EC2 and NLB control planes into 1,000+ downstream internet services worldwide. Duration substantial (~14.5h to full recovery) though the root DNS fault was corrected within ~2.6h. User and financial scores reflect millions of Downdetector reports and >$1B press-estimated economic impact, tempered because AWS disclosed no official user or dollar figures.

Sequence of events (SOE)

Phased sequence of eventspre-incident · TRIGGER — Latent race condition dormant in the DynamoDB DNS management system: a point-in-time staleness check plus unsynchronised Planner/Enactor cleanup, waiting on a rare delay-timing interleaving.TRIGGERpre-incident2025-10-19 ~11:47 PM PDT · TRIGGER — One DNS Enactor hits unusually high delays applying endpoint updates; the Planner emits many newer plan generations; a second Enactor applies the newest plan and starts cleanup, and the delayed Enactor then overwrites the regional endpoint with its stale plan (time approximate, just before the 11:48 PM error onset).TRIGGER2025-10-19 ~11:47 PM PD2025-10-19 11:48 PM PDT · IMPACT — Cleanup deletes the stale plan; all IPs for the DynamoDB US-EAST-1 regional endpoint are removed. DynamoDB API errors begin; customers cannot establish new connections. EC2 DWFM state checks (DynamoDB-dependent) begin failing.IMPACT2025-10-19 11:48 PM PD2025-10-20 12:11 AM PDT · CASCADE — Beginning at 12:11 AM PDT customers experienced increased error rates and latencies across multiple AWS services in US-EAST-1 as the foundational dependency propagated.CASCADE2025-10-20 12:11 AM PD2025-10-20 12:38 AM PDT · DETECTION — Engineers identify DynamoDB's DNS state as the source of the outage, roughly 50 minutes after onset.DETECTION2025-10-20 12:38 AM PD2025-10-20 ~12:38-1:15 AM PDT · MITIGATION — No automated self-repair activates: the empty/inconsistent record 'prevented subsequent plan updates from being applied by any DNS Enactors,' so recovery must be fully manual.MITIGATION2025-10-20 ~12:38-1:15 AM PD2025-10-20 1:15 AM PDT · MITIGATION — Temporary mitigations enable internal services to reconnect and recovery tooling to run; operators begin manual reconstruction of the DNS state.MITIGATION2025-10-20 1:15 AM PD2025-10-20 2:24 AM PDT · CASCADE — EC2 droplet leases time out as DWFM remains unable to reach DynamoDB.CASCADE2025-10-20 2:24 AM PD2025-10-20 2:25 AM PDT · RECOVERY — All DNS information restored; DynamoDB endpoints begin resolving. DWFM immediately enters congestive collapse attempting mass lease re-establishment.RECOVERY2025-10-20 2:25 AM PD2025-10-20 2:32 AM PDT · RECOVERY — All DynamoDB global-tables replicas fully caught up.RECOVERY2025-10-20 2:32 AM PD2025-10-20 2:25-2:40 AM PDT · RECOVERY — Customer endpoint resolution completes as cached empty DNS records expire from resolver caches; DynamoDB recovery complete at 2:40 AM.RECOVERY2025-10-20 2:25-2:40 AM PD2025-10-20 4:14 AM PDT · MITIGATION — Engineers throttle incoming EC2 work and selectively restart DWFM hosts to break congestive collapse and clear queues.MITIGATION2025-10-20 4:14 AM PD2025-10-20 5:28 AM PDT · RECOVERY — EC2 droplet leases re-established; Network Manager begins working through a large network-state propagation backlog.RECOVERY2025-10-20 5:28 AM PD2025-10-20 9:36 AM PDT · MITIGATION — Engineers disable NLB automatic health-check failovers to stop capacity being repeatedly removed by flapping instances, keeping healthy capacity online.MITIGATION2025-10-20 9:36 AM PD2025-10-20 10:36 AM PDT · RECOVERY — New EC2 instance connectivity restored as the network-state propagation backlog drains.RECOVERY2025-10-20 10:36 AM PD2025-10-20 1:50 PM PDT · RECOVERY — EC2 incoming-work throttles fully removed; full EC2 recovery achieved as DWFM stabilises.RECOVERY2025-10-20 1:50 PM PD2025-10-20 2:09 PM PDT · RESTORED — NLB automatic health-check failovers re-enabled after EC2 stabilises; overall event concluded ~2:20 PM. AWS separately disables the DynamoDB DNS Planner and Enactor automation worldwide pending fixes.RESTORED2025-10-20 2:09 PM PD

Root cause

The immediate trigger of the October 20, 2025 disruption was a latent race condition in Amazon DynamoDB's automated DNS management system in the US-EAST-1 (Northern Virginia) Region. Per the AWS post-event summary, this race condition resulted in "an incorrect empty DNS record for the service's regional endpoint (dynamodb.us-east-1.amazonaws.com)," and the automation was then unable to repair that state. DynamoDB's endpoint DNS is managed by two cooperating automation components: a DNS Planner, which "monitors the health and capacity of the load balancers and periodically creates a new DNS plan," and a DNS Enactor, which applies those plans into Amazon Route53 and which is "designed to have minimal dependencies to allow for system recovery," operating "redundantly and fully independently in three different Availability Zones (AZs)." The mechanism of failure was a timing collision between two independently operating Enactors. Per AWS, "Before it begins to apply a new plan, the DNS Enactor makes a one-time check that its plan is newer than the previously applied plan." On this occasion one Enactor experienced "unusually high delays" while retrying endpoint updates as the Planner produced newer plan generations. Meanwhile a second Enactor rapidly applied the newest plan across all endpoints and then "invoked the plan clean-up process, which identifies plans that are significantly older than the one it just applied and deletes them." The delayed first Enactor then applied "its much older plan to the regional DDB endpoint, overwriting the newer plan" — its one-time freshness check "was stale by this time due to the unusually high delays." When the clean-up process deleted that just-applied older plan, "all IP addresses for the regional endpoint were immediately removed," leaving an empty DNS record. The safeguards did not prevent the outcome. The redundancy design — three fully independent Enactors across three AZs — was intended to survive AZ or host failures, but this was not an infrastructure failure; it was a logical race between the redundant actors themselves, independently mutating shared Route53 state with coordination resting only on a one-time staleness check that could go stale under processing delay, paired with an aggressive clean-up step that deleted the very plan the delayed Enactor had just written. Worse, once the record was emptied, "the system was left in an inconsistent state that prevented subsequent plan updates from being applied by any DNS Enactors," so the automation could not self-heal and manual operator intervention was required. Escalation was immediate and broad. Systems connecting to DynamoDB in us-east-1 via the public endpoint failed DNS resolution and could not connect. The failure cascaded into dependent services. EC2's DropletWorkflow Manager (DWFM), which maintains a lease for each managed droplet, began failing state checks: "Starting at 11:48 PM PDT on October 19, these DWFM state checks began to fail as the process depends on DynamoDB and was unable to complete." Even after DynamoDB DNS was restored, "DWFM had entered a state of congestive collapse and was unable to make forward progress in recovering droplet leases," requiring engineers to throttle incoming work (at 4:14 AM PDT) and perform selective host restarts (leases re-established by 5:28 AM). DWFM recovery then produced "a significant backlog of network state propagations" in Network Manager, delaying connectivity for newly launched EC2 instances. During that window, propagation lag caused NLB health checks to fail against otherwise healthy nodes: per AWS, "health checks would fail even though the underlying NLB node and backend targets were healthy," which "resulted in health checks alternating between failing and healthy" and "caused NLB nodes and backend targets to be removed from DNS, only to be returned to service when the next health check succeeded" — producing further customer-visible errors. Detection was fast: DynamoDB DNS was identified as the source at 12:38 AM PDT (roughly 50 minutes after the 11:48 PM onset); all DNS information was manually restored by 2:25 AM and primary DynamoDB disruption ended by 2:40 AM. The downstream EC2/Network Manager/NLB cascade extended recovery, with all EC2 APIs and new instance launches operating normally by 1:50 PM PDT (13:50).

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What happened

A latent race condition in DynamoDB's automated DNS management for US-EAST-1 produced an empty DNS record for the service's regional endpoint. A delayed DNS Enactor overwrote a current plan with a stale one, and the parallel cleanup then deleted that plan, removing all IP addresses. AWS: 'The root cause of this issue was a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record for the service's regional endpoint' (aws.amazon.com/message/101925/).

Why the guard failed

The Enactor's staleness check ran only at the start of plan application and 'was stale by this time due to the unusually high delays,' so it could not stop a superseded plan from overwriting the live record. Three Enactors across three AZs shared no apply-time mutual exclusion, so cleanup could delete a plan another Enactor was concurrently writing. The result was a self-inflicted empty record the automation could not self-heal (aws.amazon.com/message/101925/).

How it cascaded

DynamoDB is a foundational dependency. EC2's DWFM lost DynamoDB access, droplet leases timed out, and on DNS restoration DWFM entered 'congestive collapse' attempting mass lease re-establishment. The NLB health-check subsystem flapped on newly launched instances and triggered automatic AZ failover that stripped capacity. Downstream, Roblox, Fortnite, Snapchat and Duolingo among others degraded (aws.amazon.com/message/101925/; en.wikipedia.org).

Detection and response

Onset at 11:48 PM PDT; correct diagnosis by 12:38 AM (~50 min). Because no automated self-repair existed, recovery was manual: DNS rebuilt by 2:25 AM, congestive collapse broken at 4:14 AM by throttling and host restarts, new-EC2 connectivity by 10:36 AM, NLB failover re-enabled 2:09 PM, event concluded ~2:20 PM. The DWFM scenario had 'no established operational recovery procedure' (aws.amazon.com/message/101925/).

Systemic lessons and fixes

The remediation set is an implicit admission of latent design and test gaps: fix the race and add stale-plan protection, add NLB velocity control, build a DWFM recovery test suite, add queue-depth-based EC2 throttling, and disable the DNS automation worldwide until fixes ship. The transferable lessons: validate invariants at write time, coordinate redundant writers, test recovery paths, and add backpressure so recovery cannot overload the system it is recovering (aws.amazon.com/message/101925/).

Technical deep-dive

DynamoDB's regional endpoint in US-EAST-1 is resolved via DNS records maintained by an internal control loop split across two roles for availability. The DNS Planner watches load-balancer health and capacity and emits monotonically increasing "plans" (generations) describing the correct set of endpoints, load balancers and weights. The DNS Enactor is the applier — running independently in each of three Availability Zones, each pushing the latest plan to DNS and then running a cleanup pass that garbage-collects plans many generations old. Redundancy here was for availability, but the three Enactors shared no apply-time mutual exclusion or compare-and-swap against the live record, so redundancy became a hazard multiplier rather than a safety net. The bug is a classic time-of-check-to-time-of-use (TOCTOU) window. Enactor A reads plan generation N, passes the "newer-than-previously-applied" check, then stalls under unusually high delays. Meanwhile the Planner emits N+1, N+2, ... and Enactor B applies the newest generation and triggers cleanup, deleting older generations. Enactor A, resuming with its now-stale in-hand plan, writes it over the current record — the guard cannot catch this because it was evaluated before the stall, not re-evaluated at write time. Cleanup then deletes the plan Enactor A just re-applied because it is generations behind, and the endpoint record collapses to empty. An empty DNS record is worse than a stale one: clients cannot resolve the endpoint at all, region-wide, and because the active plan was deleted the automation refuses to apply any further plan, seeing an inconsistent state — the self-heal path is bricked, forcing manual operator repair. Because DynamoDB is a foundational dependency, the empty record cascaded into control planes. EC2's DropletWorkflow Manager (DWFM) maintains droplet leases tracking physical-server state and depends on DynamoDB; lease health checks began failing during the outage. When DNS was restored at 2:25 AM PDT, DWFM tried to re-establish every lease at once but could not complete reconnections before timeouts triggered retries, and "DWFM had entered a state of congestive collapse and was unable to make forward progress in recovering droplet leases" — a positive-feedback overload where recovery work itself prevents recovery. Because "this situation had no established operational recovery procedure," engineers proceeded cautiously; at 4:14 AM they throttled incoming work and selectively restarted DWFM hosts, with leases re-established by 5:28 AM. New-instance connectivity lagged until 10:36 AM because Network Manager had to drain a large network-state propagation backlog. In parallel, the Network Load Balancer health-check subsystem saw newly launched instances (arriving before network state fully propagated) flap between healthy and failing, driving automatic AZ DNS failover that repeatedly removed and restored capacity; engineers disabled automatic health-check failover at 9:36 AM (re-enabled 2:09 PM) to arrest the churn. Caching shaped the tail: even after the record was rebuilt at 2:25 AM, customer resolution only fully cleared by 2:40 AM as cached empty answers expired. Downstream, consumer-facing services including Roblox, Fortnite, Snapchat and Duolingo saw elevated errors and latency. All times per aws.amazon.com/message/101925/ except downstream service names (en.wikipedia.org).

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home