← All incidents
Incident dossier · Rank #28

AT&T Nationwide Wireless Outage — Network Misconfiguration (February 22, 2024)

AT&T 2024-02-22 12h 0m core impact NetworkSoftwareHuman error

During a network expansion, an AT&T engineer used an incorrect process that introduced a misconfiguration into the wireless network, causing a roughly 12-hour nationwide outage affecting an estimated 125 million devices and blocking more than 92,000 emergency 911 calls.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2024-02-22)Trigger · Network2024-02-222024-02-22Primary fault at AT&T — AT&T national wireless core networkAT&T MobilityAT&T national wireless core networkAT&T national wireless coreDownstream service degraded by the fault: AT&T wireless voice/data nationwideAT&T wireless voice/dataDownstream service degraded by the fault: 911 emergency calling (partial)911 emergency calling

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
AT&T
Data center
AT&T national wireless core network
Location
Dallas, Texas, United States
Date
2024-02-22

Impact & scale

Users affected
~125 million devices nationwide; 92,000+ blocked 911 calls
Financial
FCC penalties + customer credits
Scope
Sev-1 national wireless core (config)
Services / systems down
  • AT&T wireless voice/data nationwide
  • 911 emergency calling (partial)

Impact data & metrics

Time from misconfigured element insertion to nationwide protective shutdown3 minutes (02:42 AM to 02:45 AM CT)
Total outage durationAt least twelve hours
Registered devices affectedMore than 125 million
Voice calls blockedMore than 92 million
911/PSAP call attempts blockedMore than 25,000
Time to roll back the network changeClose to 2 hours
FirstNet infrastructure restoration completedBy 05:00 AM CT
Delay before FirstNet customers were notifiedNotification began 05:53 AM, more than 3 hours after outage began
Geographic scopeAll 50 states, DC, Puerto Rico, and US Virgin Islands
Wireless Priority Service call volume during outageLess than two-thirds of typical volume
Time to fully remediate registration congestion after restorationMore than 10 hours
Post-outage window to scan network and add missing mitigating controlsWithin 48 hours

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 5Financial impact (0–10) — direct + consequential cost. — scored 5/10.Duration 5Outage duration (0–10) — how long service was degraded/down. — scored 5/10.Blast 8Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 8/10.
Magnitude 6.8 = blast 8×0.35 + users 8×0.25 + financial 5×0.20 + duration 5×0.20 (sub-scores 0–10 · weighted composite)

A network misconfiguration during a capacity expansion caused a ~12h nationwide wireless outage, ~125M devices, tens of thousands of blocked 911 calls — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events2024-02-22 02:42 AM CT · TRIGGER — An AT&T Mobility employee places a new, misconfigured network element into the production wireless core during a routine night maintenance window to expand capacity; the configuration does not conform to procedures requiring peer review, and that peer review was never performed.TRIGGER2024-02-22 02:42 AM C~02:42-02:45 AM CT · CASCADE — 'As a result of the error in configuration, downstream network elements propagated the error further into the network' — the fault spreads from the single new element to downstream elements within ~3 minutes.CASCADE~02:42-02:45 AM C~02:42-02:45 AM CT · DETECTION — Detection is AUTOMATED and effectively instantaneous, indistinguishable from the protective response: the propagating error 'triggered an automated response that shut down all network connections.' The report describes no separate alarm/telemetry step and states no human-detection time.DETECTION~02:42-02:45 AM C2024-02-22 02:45 AM CT · MITIGATION — Automated 'Protection Mode' trip fires 'just three minutes after the misconfigured network element was placed into production' — the designed containment (fire-suppression analogue) activates as intended to stop error propagation.MITIGATION2024-02-22 02:45 AM C2024-02-22 02:45 AM CT · IMPACT — Protection Mode is over-broad: the shutdown 'isolated all voice and 5G data processing elements from the wireless towers and switching elements,' disconnecting all registered devices — including FirstNet — from voice and 5G data, beginning a nationwide outage.IMPACT2024-02-22 02:45 AM C2024-02-22 from 02:45 AM CT · IMPACT — Outage affects more than 125 million registered devices across all 50 states, DC, Puerto Rico and the US Virgin Islands, including AT&T MVNO and roaming subscribers on AT&T's network.IMPACT2024-02-22 from 02:45 AM C2024-02-22 during outage · IMPACT — Emergency services crippled: no 911 calls from AT&T Mobility-served wireless devices could be routed to the destination PSAP while voice services were disconnected; more than 25,000 calls to PSAPs/911 were blocked.IMPACT2024-02-22 durin2024-02-22 during outage · IMPACT — Priority public-safety calling degraded: Wireless Priority Service (WPS) call volume ran at 'less than two thirds of the typical volume,' indicating some priority calls also failed.IMPACT2024-02-22 durin2024-02-22 (starts ~03:xx AM CT) · RECOVERY — AT&T performs a rollback that backs out / removes the misconfigured network element and begins restoration of the network — the corrective isolation step (removing the faulty element).RECOVERY2024-02-22 (starts ~03:xx AM C~2 hours · RECOVERY — 'It took close to two hours to roll back the network change' before normal-operations restoration could proceed.RECOVERY~2 hours2024-02-22 05:00 AM CT · RECOVERY — 'By 5:00 AM on February 22, 2024, FirstNet infrastructure was restored' — public-safety infrastructure prioritized.RECOVERY2024-02-22 05:00 AM C2024-02-22 after rollback · CASCADE — Re-registration STORM (second-order failure): once the element was backed out, 'all user devices automatically tried to simultaneously re-register'; 'The influx of device registration attempts far exceeded what AT&T Mobility's network management systems could handle, resulting in widespread congestion,' prolonging the outage.CASCADE2024-02-22 after2024-02-22 05:53 AM CT · MITIGATION — FirstNet-customer notification begins, 'more than three hours after the outage began' and nearly one hour after FirstNet service was restored — the evacuation/notification analogue, sent late.MITIGATION2024-02-22 05:53 AM C2024-02-22 07:05 AM CT · MITIGATION — AT&T Mobility issues its first public statement about the outage.MITIGATION2024-02-22 07:05 AM C2024-02-22 06:46 PM CT · RESTORED — AT&T Mobility publicly states initial findings: the outage 'was caused by the application and execution of an incorrect process' during a capacity expansion, and was 'not a cyberattack.'RESTORED2024-02-22 06:46 PM C2024-02-22 (~12h+ from 02:45 AM) · RESTORED — Full voice and 5G data service restored after an outage lasting at least twelve hours; the FCC PSHSB later referred the matter to the Enforcement Bureau for potential violations of Parts 4 and 9 of the Commission's rules.RESTORED2024-02-22 (~12h

Root cause

The proximate trigger was not combustion but a NEW NETWORK ELEMENT placed into AT&T Mobility's production wireless core at 2:42 AM CT on February 22, 2024, during a routine night maintenance window to expand network capacity (FCC DOC-404150A1, para 6). That element was misconfigured. Its configuration "did not conform to AT&T's established network element design and installment procedures, which require peer review" (para 6). MAKE / MODEL / VENDOR / ELEMENT-TYPE / AGE ARE NOT DISCLOSED: the FCC report deliberately describes only a generic "network element" and "downstream network elements," names no vendor, hardware/software make or model, or core function, and — being a software/configuration fault, not a physical one — there is no "chemistry." AT&T's own public account attributed the event to "the application and execution of an incorrect process" during a capacity expansion, and explicitly "not a cyberattack" (AT&T 6:46 PM statement quoted at para 11); AT&T's separate public wording of a "coding error" is consistent paraphrase, not the FCC's language. The failure mechanism (the ignition-to-fire chain) is a four-step cascade: (1) the non-conforming configuration was loaded without the required peer review; (2) "As a result of the error in configuration, downstream network elements propagated the error further into the network" (para 6); (3) this "triggered an automated response that shut down all network connections to prevent the traffic from propagating further" — the designed "Protection Mode" trip; and (4) that shutdown "isolated all voice and 5G data processing elements from the wireless towers and switching elements," disconnecting registered devices nationwide at 2:45 AM, "just three minutes after the misconfigured network element was placed into production" (para 6). Critically, "the downstream network element lacked controls specific to mitigating this error and therefore was unable to mitigate the effects" (para 25) — there was no granular quarantine, so the only available containment was a nationwide disconnect. The LATENT ROOT is a stack of design, testing and process (maintenance) failures. Design: Protection Mode was engineered as all-or-nothing containment with no ability to isolate a single bad element (para 25), and the network had no plan for the device re-registration storm that a Protection-Mode recovery guarantees (para 27). Testing/maintenance: "Lab testing by AT&T did not discover the improper configuration of the network element that caused the outage and did not identify the potential impact" (para 22), and post-installation testing on the night of Feb 22 was inadequate to identify the incorrect behavior, contrary to CSRIC Best Practice 13-10-0615 (para 21). Process: the mandatory peer review did not take place, and "Adequate peer review should have prevented the network change from being approved, and, in turn, from being loaded onto the network" (para 20). The single latent root that most directly caused the event is therefore the failure to perform — and to enforce as a gate — the required peer review of a production change to the wireless core.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What happened: an unreviewed change, a three-minute cascade

At 2:42 AM CT on February 22, 2024, an AT&T Mobility employee inserted a new, misconfigured network element into the production wireless core during a routine maintenance window intended to expand capacity. The configuration violated AT&T's design-and-installment procedures, which require peer review — a review that never occurred. The error propagated to downstream elements and, at 2:45 AM, triggered an automated shutdown, three minutes after insertion. No vendor, make, model, or element function is disclosed by the FCC; the fault was software/configuration, not physical, so there is no hardware chemistry to describe (DOC-404150A1, para 6, 20).

Why one bad element became a nationwide outage: Protection Mode

The decisive design factor is that the downstream element 'lacked controls specific to mitigating this error,' so the network had no way to quarantine a single misconfigured element. Instead it invoked 'Protection Mode,' an automated total shutdown that 'isolated all voice and 5G data processing elements from the wireless towers and switching elements.' Protection Mode worked as designed — it stopped error propagation — but its blast radius was the entire network: every registered device across all 50 states, DC, Puerto Rico and the US Virgin Islands lost voice and 5G data at once (para 25, 6, 1).

Public-safety impact: 911, PSAPs, FirstNet and WPS

The outage disconnected public-safety users along with everyone else. No 911 call from an AT&T Mobility device could be routed to its PSAP while voice was down, and more than 25,000 calls to PSAPs/911 were blocked. FirstNet devices were affected; FirstNet infrastructure was prioritized and restored by 5:00 AM, but FirstNet customers were not notified until 5:53 AM — more than three hours after the outage began. Wireless Priority Service volume ran at less than two-thirds of typical, indicating some priority public-safety calls also failed (para 1, 8, 9, 14, 17).

The recovery storm: a second failure designed in

Removing the misconfigured element was necessary but insufficient. Once it was backed out, all user devices simultaneously tried to re-register, and 'the influx of device registration attempts far exceeded what AT&T Mobility's network management systems could handle,' causing widespread congestion that prolonged the outage. AT&T 'failed to prepare for the registration congestion associated with the network recovering from Protection Mode,' and it took more than 10 hours after restoration to fully remediate the congestion — making the recovery, not the initial trip, the dominant contributor to the roughly twelve-hour total (para 26, 27).

Regulatory response and remedies

The FCC Public Safety and Homeland Security Bureau opened its investigation and produced this report, referring the matter to the Enforcement Bureau for potential violations of Parts 4 and 9 of the Commission's rules. AT&T's own remedies were concrete: within 48 hours it scanned the network for elements lacking mitigating controls and added them, and it implemented additional peer-review steps plus procedures ensuring maintenance cannot proceed without confirmation that required peer reviews are complete. The report also faults inadequate lab and post-installation testing against CSRIC Best Practice 13-10-0615 (para 21, 22, 28, 29).

Technical deep-dive

This was a "sunny day" outage — no weather, no attack, no hardware fire — in which an automated protective mechanism, functioning exactly as designed, converted a single misconfigured element into a nationwide loss of service. At 2:42 AM CT an AT&T Mobility employee placed a new, misconfigured network element into the production core during a routine night maintenance window (DOC-404150A1, para 6). The misconfiguration violated AT&T's design-and-installment procedures, which require peer review, and that peer review was never performed (para 6, 20). Within moments the error propagated to downstream elements, which "lacked controls specific to mitigating this error" (para 25) and instead invoked "Protection Mode" — an automated total shutdown that isolates processing elements to stop error propagation. Protection Mode is the network analogue of a fire-suppression system: it activated as designed and DID contain the fault (no wider corruption), but it was catastrophically over-broad, "isolat[ing] all voice and 5G data processing elements from the wireless towers and switching elements" and disconnecting every device — including FirstNet public-safety devices — across all 50 states, DC, Puerto Rico and the US Virgin Islands at 2:45 AM, three minutes after insertion (para 1, 6). Detection was automated and effectively instantaneous, indistinguishable from the protective trip itself; the report describes no separate alarm/telemetry step and gives no human-detection time. Recovery required a manual rollback that "took close to two hours" (para 2); FirstNet infrastructure was prioritized and restored by 5:00 AM (para 8). Lifting Protection Mode then produced a second-order failure: "once the misconfigured network element was backed out, all user devices automatically tried to simultaneously re-register," and "The influx of device registration attempts far exceeded what AT&T Mobility's network management systems could handle, resulting in widespread congestion" that "caused devices to be delayed in registering to the network, thereby prolonging the outage" (para 26). AT&T "failed to prepare for the registration congestion associated with the network recovering from Protection Mode, or to sufficiently mitigate that congestion after the fact," and it took "more than 10 hours once the network was restored to fully remediate the network congestion" (para 27). The regulatory record is emphatic on public-safety impact: more than 25,000 calls to PSAPs/911 were blocked (para 1), no 911 call from an AT&T-served device could be routed to its PSAP while voice was down (para 14), and Wireless Priority Service volume ran at "less than two thirds of the typical volume," implying some priority public-safety calls also failed (para 17). The FCC's Public Safety and Homeland Security Bureau referred the matter to the Enforcement Bureau for potential violations of Parts 4 and 9 of the Commission's rules (para 29).

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home