AT&T Nationwide Wireless Outage — Network Misconfiguration (February 22, 2024)
During a network expansion, an AT&T engineer used an incorrect process that introduced a misconfiguration into the wireless network, causing a roughly 12-hour nationwide outage affecting an estimated 125 million devices and blocking more than 92,000 emergency 911 calls.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- AT&T
- Data center
- AT&T national wireless core network
- Location
- Dallas, Texas, United States
- Date
- 2024-02-22
Impact & scale
- Users affected
- ~125 million devices nationwide; 92,000+ blocked 911 calls
- Financial
- FCC penalties + customer credits
- Scope
- Sev-1 national wireless core (config)
- AT&T wireless voice/data nationwide
- 911 emergency calling (partial)
Impact data & metrics
| Time from misconfigured element insertion to nationwide protective shutdown | 3 minutes (02:42 AM to 02:45 AM CT) |
| Total outage duration | At least twelve hours |
| Registered devices affected | More than 125 million |
| Voice calls blocked | More than 92 million |
| 911/PSAP call attempts blocked | More than 25,000 |
| Time to roll back the network change | Close to 2 hours |
| FirstNet infrastructure restoration completed | By 05:00 AM CT |
| Delay before FirstNet customers were notified | Notification began 05:53 AM, more than 3 hours after outage began |
| Geographic scope | All 50 states, DC, Puerto Rico, and US Virgin Islands |
| Wireless Priority Service call volume during outage | Less than two-thirds of typical volume |
| Time to fully remediate registration congestion after restoration | More than 10 hours |
| Post-outage window to scan network and add missing mitigating controls | Within 48 hours |
Magnitude profile
A network misconfiguration during a capacity expansion caused a ~12h nationwide wireless outage, ~125M devices, tens of thousands of blocked 911 calls — sub-scores ESTIMATED from public impact reporting, pending deep research.
Sequence of events (SOE)
- TRIGGER An AT&T Mobility employee places a new, misconfigured network element into the production wireless core during a routine night maintenance window to expand capacity; the configuration does not conform to procedures requiring peer review, and that peer review was never performed.
- CASCADE 'As a result of the error in configuration, downstream network elements propagated the error further into the network' — the fault spreads from the single new element to downstream elements within ~3 minutes.
- DETECTION Detection is AUTOMATED and effectively instantaneous, indistinguishable from the protective response: the propagating error 'triggered an automated response that shut down all network connections.' The report describes no separate alarm/telemetry step and states no human-detection time.
- MITIGATION Automated 'Protection Mode' trip fires 'just three minutes after the misconfigured network element was placed into production' — the designed containment (fire-suppression analogue) activates as intended to stop error propagation.
- IMPACT Protection Mode is over-broad: the shutdown 'isolated all voice and 5G data processing elements from the wireless towers and switching elements,' disconnecting all registered devices — including FirstNet — from voice and 5G data, beginning a nationwide outage.
- IMPACT Outage affects more than 125 million registered devices across all 50 states, DC, Puerto Rico and the US Virgin Islands, including AT&T MVNO and roaming subscribers on AT&T's network.
- IMPACT Emergency services crippled: no 911 calls from AT&T Mobility-served wireless devices could be routed to the destination PSAP while voice services were disconnected; more than 25,000 calls to PSAPs/911 were blocked.
- IMPACT Priority public-safety calling degraded: Wireless Priority Service (WPS) call volume ran at 'less than two thirds of the typical volume,' indicating some priority calls also failed.
- RECOVERY AT&T performs a rollback that backs out / removes the misconfigured network element and begins restoration of the network — the corrective isolation step (removing the faulty element).
- RECOVERY 'It took close to two hours to roll back the network change' before normal-operations restoration could proceed.
- RECOVERY 'By 5:00 AM on February 22, 2024, FirstNet infrastructure was restored' — public-safety infrastructure prioritized.
- CASCADE Re-registration STORM (second-order failure): once the element was backed out, 'all user devices automatically tried to simultaneously re-register'; 'The influx of device registration attempts far exceeded what AT&T Mobility's network management systems could handle, resulting in widespread congestion,' prolonging the outage.
- MITIGATION FirstNet-customer notification begins, 'more than three hours after the outage began' and nearly one hour after FirstNet service was restored — the evacuation/notification analogue, sent late.
- MITIGATION AT&T Mobility issues its first public statement about the outage.
- RESTORED AT&T Mobility publicly states initial findings: the outage 'was caused by the application and execution of an incorrect process' during a capacity expansion, and was 'not a cyberattack.'
- RESTORED Full voice and 5G data service restored after an outage lasting at least twelve hours; the FCC PSHSB later referred the matter to the Enforcement Bureau for potential violations of Parts 4 and 9 of the Commission's rules.
Root cause
Contributing factors
- PROCESS / peer-review gap (central lapse): the change violated procedures that require peer review and that peer review did not take place; 'Adequate peer review should have prevented the network change from being approved, and, in turn, from being loaded onto the network' (DOC-404150A1, para 6, 20).
- MAINTENANCE / inadequate lab pre-deployment testing: 'Lab testing by AT&T did not discover the improper configuration of the network element that caused the outage and did not identify the potential impact to the network of that or similar misconfigurations' (para 22).
- MAINTENANCE / inadequate post-installation testing during the maintenance window: post-installation testing is described as a well-established best practice (CSRIC Best Practice 13-10-0615) to verify complex changes and test after committing, and testing performed that night was inadequate to identify the incorrect behavior (para 21).
- DESIGN / missing granular mitigating controls: 'The downstream network element lacked controls specific to mitigating this error and therefore was unable to mitigate the effects' — forcing a nationwide protective shutdown instead of a quarantine (para 25).
- DESIGN + PROCESS / no recovery-capacity planning for the re-registration storm: 'AT&T failed to prepare for the registration congestion associated with the network recovering from Protection Mode, or to sufficiently mitigate that congestion after the fact'; remediation took more than 10 hours after restoration (para 26, 27).
- PROCESS / weak change-approval controls allowed the change onto the production network without confirmation that required reviews were completed, later remediated by procedures ensuring maintenance cannot proceed without confirmed peer review (para 20, 28).
Correction of errors (COE)
- Implement additional peer-review steps and procedures ensuring maintenance cannot proceed without confirmation that required peer reviews are completed
- Scan the production network for elements lacking mitigating controls and add element-specific controls to enable quarantine instead of nationwide shutdown
- Strengthen lab and post-installation testing to detect misconfigurations and model network-wide impact per CSRIC Best Practice 13-10-0615
- Build higher-capacity, controlled device re-registration/recovery systems to survive a Protection-Mode recovery storm
- Prioritize 911/PSAP routing restoration and accelerate FirstNet/public-safety outage notification
- Enforcement Bureau review of potential Part 4 and Part 9 rule violations
Lessons learnt
- A single unreviewed configuration change to a shared core can become a nationwide outage in minutes — process controls (peer review) are not paperwork, they are the primary safety barrier, and they must be an enforced gate, not an advisory step (para 6, 20).
- Automated protective mechanisms must be scoped to the smallest blast radius that still contains the fault; an all-or-nothing 'Protection Mode' that isolates every processing element trades a local error for a national one (para 25).
- Design for the recovery, not just the failure: a protective shutdown guarantees a synchronized device re-registration storm, and without pre-planned capacity that storm becomes the dominant driver of outage duration (over 10 hours of remediation) (para 26, 27).
- Testing that does not reproduce production topology and does not model network-wide impact provides false assurance — lab and post-installation testing must be able to catch a misconfiguration and its downstream propagation (para 21, 22).
- Public-safety resilience is a distinct design requirement: 911/PSAP routing and FirstNet must be prioritized in restoration and public-safety users must be notified fast — a 3-hour notification lag is itself a failure (para 9, 14).
Improvements & remediation
- ProcessEnforce mandatory peer review as a hard gate on all production wireless-core changes — post-outage AT&T 'implemented additional steps for peer review' and 'adopted procedures to ensure that maintenance work cannot take place without confirmation that required peer reviews have been completed' (DOC-404150A1, para 28).
- MaintenanceStrengthen lab and post-installation testing to detect misconfigurations and model their network-wide impact before and immediately after a change, per CSRIC Best Practice 13-10-0615 — verify complex configuration changes before committing and test after, closing the gap where 'Lab testing by AT&T did not discover the improper configuration' (para 21-22).
- DesignAdd element-specific mitigating controls so a single misconfigured element can be quarantined rather than triggering a nationwide Protection-Mode shutdown; within 48 hours AT&T scanned the network for elements lacking the controls and put those controls in place (para 25, 28).
- DesignEngineer for the recovery, not just the failure — build re-registration congestion controls and greater registration capacity so devices re-attach in a controlled sequence after Protection Mode lifts, since 'More robust registration systems with greater capacity would have enabled AT&T Mobility to more quickly and efficiently recover' (para 27).
- SafetyGuarantee 911/FirstNet resilience and rapid public-safety notification — prioritize PSAP-routing restoration and notify FirstNet/public-safety users promptly (the 3-hour notification delay must be closed), given 'no 911 calls from AT&T Mobility-served wireless devices could be routed to the destination PSAP' while voice was down (para 9, 14).
- ProcessTighten change-approval controls so a production change cannot be loaded without recorded confirmation that all required reviews and pre-checks were completed, since 'Adequate peer review should have prevented the network change from being approved' (para 20, 28).
Comprehensive analysis
What happened: an unreviewed change, a three-minute cascade
At 2:42 AM CT on February 22, 2024, an AT&T Mobility employee inserted a new, misconfigured network element into the production wireless core during a routine maintenance window intended to expand capacity. The configuration violated AT&T's design-and-installment procedures, which require peer review — a review that never occurred. The error propagated to downstream elements and, at 2:45 AM, triggered an automated shutdown, three minutes after insertion. No vendor, make, model, or element function is disclosed by the FCC; the fault was software/configuration, not physical, so there is no hardware chemistry to describe (DOC-404150A1, para 6, 20).
Why one bad element became a nationwide outage: Protection Mode
The decisive design factor is that the downstream element 'lacked controls specific to mitigating this error,' so the network had no way to quarantine a single misconfigured element. Instead it invoked 'Protection Mode,' an automated total shutdown that 'isolated all voice and 5G data processing elements from the wireless towers and switching elements.' Protection Mode worked as designed — it stopped error propagation — but its blast radius was the entire network: every registered device across all 50 states, DC, Puerto Rico and the US Virgin Islands lost voice and 5G data at once (para 25, 6, 1).
Public-safety impact: 911, PSAPs, FirstNet and WPS
The outage disconnected public-safety users along with everyone else. No 911 call from an AT&T Mobility device could be routed to its PSAP while voice was down, and more than 25,000 calls to PSAPs/911 were blocked. FirstNet devices were affected; FirstNet infrastructure was prioritized and restored by 5:00 AM, but FirstNet customers were not notified until 5:53 AM — more than three hours after the outage began. Wireless Priority Service volume ran at less than two-thirds of typical, indicating some priority public-safety calls also failed (para 1, 8, 9, 14, 17).
The recovery storm: a second failure designed in
Removing the misconfigured element was necessary but insufficient. Once it was backed out, all user devices simultaneously tried to re-register, and 'the influx of device registration attempts far exceeded what AT&T Mobility's network management systems could handle,' causing widespread congestion that prolonged the outage. AT&T 'failed to prepare for the registration congestion associated with the network recovering from Protection Mode,' and it took more than 10 hours after restoration to fully remediate the congestion — making the recovery, not the initial trip, the dominant contributor to the roughly twelve-hour total (para 26, 27).
Regulatory response and remedies
The FCC Public Safety and Homeland Security Bureau opened its investigation and produced this report, referring the matter to the Enforcement Bureau for potential violations of Parts 4 and 9 of the Commission's rules. AT&T's own remedies were concrete: within 48 hours it scanned the network for elements lacking mitigating controls and added them, and it implemented additional peer-review steps plus procedures ensuring maintenance cannot proceed without confirmation that required peer reviews are complete. The report also faults inadequate lab and post-installation testing against CSRIC Best Practice 13-10-0615 (para 21, 22, 28, 29).
Technical deep-dive
References & provenance
- regulatory FCC Public Safety and Homeland Security Bureau — Report on the February 22, 2024 AT&T Mobility Network Outage (DOC-404150A1)“an AT&T Mobility employee placed a new network element into its production network during a routine night maintenance window ... The configuration of the network element did not conform to AT&T's established network element design and installment procedures, which require peer review.”https://docs.fcc.gov/public/attachments/DOC-404150A1.txt
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).