← All incidents
Incident dossier · Rank #7

AWS US-EAST-1 Internal Network Congestion Outage (December 7, 2021)

Amazon Web Services 2021-12-07 6h 52m core impact NetworkSoftware

On December 7, 2021, an automated capacity-scaling operation on a service hosted in AWS's main network in US-EAST-1 provoked unexpected behavior from a large population of internal clients, generating a surge of connection activity that overwhelmed the networking devices bridging AWS's internal network and its main network. Rising latency and errors drove clients to retry more aggressively, creating a self-reinforcing congestion loop — effectively a self-inflicted internal denial of service. Because the internal network hosts foundational services (monitoring, internal DNS, authorization, and parts of the EC2 control plane), the congestion cascaded into control-plane, authentication, DNS, and API failures across dozens of services for roughly seven hours. The same congestion blinded AWS's own real-time monitoring, slowing diagnosis, and impaired the Service Health Dashboard tooling from failing over to its standby region, delaying public communication.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2021-12-07)Trigger · Network2021-12-072021-12-07Primary fault at Amazon Web Services — AWS US-EAST-1 (Northern Virginia)AWSAWS US-EAST-1 (Northern Virginia)AWS US-EAST-1Downstream service degraded by the fault: EC2 control-plane APIs (launch/describe)EC2 control-plane APIsDownstream service degraded by the fault: Route 53 APIsRoute 53 APIsDownstream service degraded by the fault: STSSTSDownstream service degraded by the fault: AWS Console / loginAWS Console / login+12 more downstream services

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Amazon Web Services
Data center
AWS US-EAST-1 (Northern Virginia)
Location
Ashburn, USA, us-east-1
Date
2021-12-07

Impact & scale

Users affected
Thousands of AWS-hosted services and their end users; consumer-facing names including Netflix, Disney+, Robinhood, and Ring were reported degraded (per press coverage, not the AWS postmortem)
Financial
Not quantified against a primary source in the record; no confirmed aggregate dollar-loss figure
Scope
SEV-1 / major multi-service regional impairment
Services / systems down
  • EC2 control-plane APIs (launch/describe)
  • Route 53 APIs
  • STS
  • AWS Console / login
  • API Gateway
  • EventBridge
  • Amazon Connect
  • ECS/EKS/Fargate APIs
  • RDS (unable to create resources — blocked EC2 launches)
  • EMR (unable to create resources — blocked EC2 launches)
  • WorkSpaces (unable to create resources — blocked EC2 launches)
  • Elastic Load Balancing APIs
  • Redshift (via STS login failures)
  • CloudWatch (monitoring delays/data loss)
  • S3/DynamoDB access via VPC endpoints
  • Support-case creation

Impact data & metrics

Trigger time07:30 AM PST, 7 Dec 2021
Total network-device impairment~6 h 52 min (07:30 -> 14:22 PST full recovery)
Route 53 API impairment window7 h 00 min (07:30 AM -> 02:30 PM PST)
EC2 API error onset07:33 AM PST (launch/describe error rates + latencies)
Support Contact Center outage (case creation)~6 h 52 min (07:33 AM -> 02:25 PM PST)
Service Health Dashboard tooling restored08:22 AM PST (~52 min after trigger; failover to standby region had failed)
Internal DNS resolution errors fully recovered09:28 AM PST (~1 h 58 min after trigger)
Congestion significantly improved01:34 PM PST
Latent code-path age'in production for many years' (previously unobserved behavior); back-off fix to be deployed over the following ~two weeks
Geographic impactMainly Eastern United States; disrupted delivery service and streaming

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 9Users affected (0–10) — breadth of the user/customer population impacted. — scored 9/10.Financial 8Financial impact (0–10) — direct + consequential cost. — scored 8/10.Duration 8Outage duration (0–10) — how long service was degraded/down. — scored 8/10.Blast 9Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 9/10.
Magnitude 8.6 = blast 9×0.35 + users 9×0.25 + financial 8×0.20 + duration 8×0.20 (sub-scores 0–10 · weighted composite)

Extreme blast radius: US-EAST-1 is AWS's largest and most-depended-upon region, and several globally-scoped control planes are single-homed there, so a device-level congestion event propagated into auth, DNS, and API failures across the ecosystem. Core duration ~6h52m (07:30 onset to full network-device recovery at 14:22 PST), with a tail to 16:41 PST for Amazon Connect; user reach was large. Financial magnitude is scored high on breadth of affected commerce but is NOT anchored to a confirmed dollar figure.

Sequence of events (SOE)

Phased sequence of events2021-12-07 07:30 AM PST · Trigger — An automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggers unexpected behavior from a large number of clients inside the internal network, producing a surge of connection activity that overwhelms the networking devices bridging the internal and main AWS networks.Trigger2021-12-07 07:30 AM PS2021-12-07 ~07:30 AM PST (concurrent; no discrete timestamp in postmortem) · Escalation — Cross-network delays raise latency and errors; clients respond with even more connection attempts and retries, forming a self-reinforcing retry storm that converts the transient surge into persistent congestion on the bridging devices. The postmortem describes this dynamic without assigning it a separate timestamp.Escalation2021-12-07 ~07:30 AM PS2021-12-07 07:33 AM PST · Impact — EC2 APIs used to launch and describe instances begin showing increased error rates and latencies; the ability to create support cases is impaired (postmortem: impaired from 7:33 AM until 2:25 PM PST).Impact2021-12-07 07:33 AM PS2021-12-07 ~07:33 AM PST (concurrent; no discrete timestamp in postmortem) · Cascade — Because foundational services (monitoring, internal DNS, authorization, and parts of the EC2 control plane) live on the congested internal network, failures spread to Route 53, STS, the AWS Console/login, API Gateway, EventBridge, ECS/EKS/Fargate, ELB, CloudWatch, and S3/DynamoDB-via-VPC-endpoint paths. RDS, EMR and WorkSpaces are impacted via inability to launch EC2 instances; Redshift via STS login failures. (Lambda operated normally.)Cascade2021-12-07 ~07:33 AM PS2021-12-07 ~07:30 AM PST (immediately at onset; no discrete timestamp) · Detection — The congestion immediately degrades real-time monitoring data for AWS's internal operations teams, impairing their ability to locate the source of congestion; the Service Health Dashboard tooling cannot fail over to its standby region, delaying public acknowledgement.Detection2021-12-07 ~07:30 AM PS2021-12-07 08:22 AM PST · Diagnosis — Service Health Dashboard update posting begins to succeed after the tooling could not fail over to standby; operators continue diagnosing with degraded telemetry and deliberately move cautiously with remediations to avoid harming the many still-healthy workloads on the main network.Diagnosis2021-12-07 08:22 AM PS2021-12-07 09:28 AM PST · Mitigation — Internal DNS is prioritized as an early recovery target; internal DNS errors fully recover, removing one of the largest sources of impact roughly two hours into the event.Mitigation2021-12-07 09:28 AM PS2021-12-07 12:35 PM PST · Mitigation — As a further mitigation, operators disable EventBridge event delivery.Mitigation2021-12-07 12:35 PM PS2021-12-07 01:15 PM PST · Recovery — As congestion improves, EC2 API error rates begin to improve (except for new instance launches, which recover later).Recovery2021-12-07 01:15 PM PS2021-12-07 01:34 PM PST · Recovery — Network congestion significantly improves.Recovery2021-12-07 01:34 PM PS2021-12-07 01:35 PM PST · Recovery — Most container-related (ECS/EKS/Fargate) API error rates return to normal; API Gateway recovery begins.Recovery2021-12-07 01:35 PM PS2021-12-07 02:22 PM PST · Restored — All network devices fully recover, ending the underlying congestion that drove the event; AWS Console access is restored (~6h52m from the 07:30 onset).Restored2021-12-07 02:22 PM PS2021-12-07 02:25 PM PST · Restored — Support-case creation capability is restored (impaired since 07:33, ~6h52m of degraded customer-support intake).Restored2021-12-07 02:25 PM PS2021-12-07 02:30 PM PST · Restored — Route 53 APIs fully recover.Restored2021-12-07 02:30 PM PS2021-12-07 02:35 PM PST · Restored — EventBridge event delivery is re-enabled.Restored2021-12-07 02:35 PM PS2021-12-07 02:40 PM PST · Restored — EC2 new instance launches fully recover.Restored2021-12-07 02:40 PM PS2021-12-07 04:28 PM PST · Restored — STS full recovery completes, ending authentication-related login failures for dependent services.Restored2021-12-07 04:28 PM PS2021-12-07 04:37 PM PST · Restored — API Gateway is largely recovered.Restored2021-12-07 04:37 PM PS2021-12-07 04:41 PM PST · Restored — Amazon Connect (dependent on API Gateway) resumes normal operations — one of the last services to recover, ~9h11m from onset.Restored2021-12-07 04:41 PM PS

Root cause

There was no fire, no equipment failure, and no physical ignition source in this event, so the conventional forensic vocabulary of make/model/chemistry does not apply; the AWS post-event summary frames a network congestive-collapse outage rather than a plant failure. The "ignition source" analog is a software/network trigger: at 7:30 AM PST on 7 December 2021, "an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network" (aws.amazon.com/message/12721/). AWS did not name the specific service being scaled, and it did not disclose the make, model, or firmware of the "networking devices between the internal network and the main AWS network" that were overwhelmed; those identifications are genuinely absent from the record. The exact failure mechanism (the "fuel plus chemistry" analog) is a latent defect in AWS's internal networking clients' back-off logic. AWS states the clients "have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event" (aws.amazon.com/message/12721/). The trigger provoked a surge of connection activity that "overwhelmed the networking devices between the internal network and the main AWS network," and because the clients would not back off, latency and errors "result[ed] in even more connection attempts and retries," producing "persistent congestion and performance issues on the devices in the internal network," a self-reinforcing congestive-collapse feedback loop (aws.amazon.com/message/12721/). The latent root is a design-and-test lapse, not maintenance in the fire-safety sense (none applies here). The defective code path was long-dormant: "This code path has been in production for many years but the automated scaling activity triggered a previously unobserved behavior" (aws.amazon.com/message/12721/). The back-off behaviour was "well tested" yet the specific trigger class was never exercised, so existing test coverage failed to catch the defect, a gap in fault-injection and scaling-scenario testing. A second, architectural root amplified the blast radius: the recovery machinery (real-time monitoring, the Service Health Dashboard failover, internal deployment systems, and the Support Contact Center) all ran on the same impaired internal network, so the tooling operators needed to diagnose and remediate was degraded by the very fault it was meant to resolve, which AWS says "further slowed our remediation efforts" (aws.amazon.com/message/12721/). The record does not disclose the specific device hardware, the exact throughput of the overwhelmed devices, or the numeric back-off parameters that failed.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What happened

At 7:30 AM PST on 7 December 2021, an automated capacity-scaling action on a service in the AWS main network triggered unexpected behavior from a large number of clients on the US-EAST-1 internal network. The resulting connection surge overwhelmed the networking devices bridging the internal and main networks, and the impairment persisted until all devices recovered at 2:22 PM PST, roughly a 6-hour-52-minute event (aws.amazon.com/message/12721/).

Why the automatic safeguards failed

The internal networking clients have 'well tested request back-off behaviors' meant to let systems recover from congestion, but 'a latent issue prevented these clients from adequately backing off during this event.' The code path 'has been in production for many years' and the scaling activity triggered a 'previously unobserved behavior,' so the damping that should have arrested the surge never engaged and the system tipped into congestive collapse (aws.amazon.com/message/12721/).

Why recovery was slow: a shared failure domain

AWS's diagnostic and recovery stack lived on the failing internal network. Real-time monitoring data became unavailable, forcing operators onto logs where they misdiagnosed the origin as internal DNS errors. Internal deployment systems were impaired, the Service Health Dashboard failed to fail over to its standby region, and the Support Contact Center could not create cases, so the very tools needed to diagnose, remediate, and communicate were degraded by the fault itself (aws.amazon.com/message/12721/).

Blast radius and what survived

The damage was concentrated in control planes, not data planes. Route 53 APIs were impaired 7:30 AM to 2:30 PM PST (no DNS changes); EC2 launch/describe APIs showed elevated errors from 7:33 AM PST; ELB APIs slowed new-load-balancer provisioning while 'existing Elastic Load Balancers remained healthy.' Already-running workloads and serving traffic largely continued; the outage bit provisioning and management operations (aws.amazon.com/message/12721/).

Prevention and open questions

AWS committed to a client back-off fix (~2 weeks), additional protective network configuration (deployed), a multi-region support system, and a new Service Health Dashboard. Open questions remain: AWS did not name the service being scaled, the make/model or throughput of the overwhelmed networking devices, or the specific back-off parameters that failed, so those details are genuinely absent rather than inferable (aws.amazon.com/message/12721/).

Technical deep-dive

This event was a control-plane congestive collapse confined to the US-EAST-1 (Northern Virginia) internal network, not a physical incident. AWS operates a bifurcated network: a "main AWS network" that hosts customer-facing services and workloads, and an "internal network" that hosts foundational services (monitoring, DNS, deployment, authentication-adjacent control functions) and communicates with the main network through a bank of intermediary networking devices. At 7:30 AM PST an automated capacity-scaling action on a service in the main network provoked "an unexpected behavior from a large number of clients inside the internal network," and the resulting "large surge of connection activity that overwhelmed the networking devices between the internal network and the main AWS network, resulting in delays for communication between these networks" (aws.amazon.com/message/12721/). Normally such a surge self-limits: the clients' "well tested request back-off behaviors" are "designed to allow our systems to recover from these sorts of congestion events." The defining pathology here is that the automatic damping mechanism did not engage: "a latent issue prevented these clients from adequately backing off during this event." With back-off suppressed, the system entered positive feedback: "These delays increased latency and errors for services communicating between these networks, resulting in even more connection attempts and retries. This led to persistent congestion and performance issues on the devices in the internal network" (aws.amazon.com/message/12721/). This is classic congestive collapse: offered load rises as goodput falls, and the devices never clear their backlog on their own. The blast radius followed the dependency graph of the internal network. Control planes degraded while data planes largely survived: "the EC2 APIs that customers use to launch new instances or to describe their current instances experienced increased error rates and latencies starting at 7:33 AM PST," and "Route 53 APIs were impaired from 7:30 AM PST until 2:30 PM PST preventing customers from making changes to their DNS entries." Crucially, running resources kept working: "Existing Elastic Load Balancers remained healthy during the event, but the elevated API error rates and latencies for the ELB APIs resulted in increased provisioning times for new load balancers" (aws.amazon.com/message/12721/). The impairment was in provisioning and control operations, not in already-serving traffic. The event was self-masking. AWS's detection and remediation stack lived on the failing network: "This congestion immediately impacted the availability of real-time monitoring data for our internal operations teams, which impaired their ability to find the source of congestion and resolve it." Blinded, operators reverted to logs and misattributed the origin, initially identifying elevated internal DNS errors and chasing that symptom, which only partly helped. Deployment tooling was likewise degraded: "our internal deployment systems, which run in our internal network, were impacted, which further slowed our remediation efforts." Even the customer escape hatches failed: the Service Health Dashboard tooling failed to fail over to the standby region (dashboard updates resumed at 8:22 AM PST) and "the ability to create support cases was impacted from 7:33 AM until 2:25 PM PST" because "our Support Contact Center also relies on the internal AWS network." Recovery was manual and staged: identifying the top sources of traffic to isolate to dedicated network devices, disabling some heavy network traffic services, and bringing additional networking capacity online, plus immediately disabling the scaling automation that triggered the event. DNS resolution errors fully recovered by 9:28 AM PST; congestion "significantly improved" by 1:34 PM PST; and by 2:22 PM PST all network devices fully recovered and console access was restored (aws.amazon.com/message/12721/). Externally, the outage "mainly affected the Eastern United States, disrupting delivery service and streaming" (en.wikipedia.org/wiki/Amazon_Web_Services).

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-07-31.

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home