Rogers Communications Canada Nationwide Network Outage (July 2022)
A maintenance change to Rogers' core IP network removed a routing filter, flooding the routers and collapsing the entire national network — taking mobile, internet, 911 emergency calling and Interac payments offline across Canada for roughly 15 hours and affecting about 12 million subscribers.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Rogers Communications
- Data center
- Rogers national IP core network
- Location
- Toronto, Canada
- Date
- 2022-07-08
Impact & scale
- Users affected
- ~12 million subscribers nationwide; 911 and Interac payment networks down
- Financial
- Regulator-ordered credits + reputational cost; exact figure not consolidated
- Scope
- Sev-1 national network core (multi-hour)
- Mobile voice/data
- Home internet
- 911 emergency calling
- Interac debit/e-transfer payments
Impact data & metrics
| Route flood magnitude (one distribution router) | Over 900,000 routes injected vs ~10,000 normally |
| Customers affected | More than 12 million (wireless + wireline) |
| Core router crash time after filter removal | Within minutes |
| Outage window (onset to full restoration) | 04:58 EDT 8 Jul → 07:00 EDT 9 Jul (~26 hours) |
| BGP update spike / peak (external detection) | Spike after 08:15 UTC, peak 08:45 UTC; prefix withdrawal from Rogers ASN |
| Traffic recovery by 09 Jul 08:40 UTC | ~76% of prior-day level at same time |
| Upgrade phase at failure | Phase 6 of a 7-phase core IP-network update |
| Reported deaths (911 unavailable) | At least one reported (avoidability unclear) |
| Wireless/wireline core separation investment | $261 million (CAD) pledged |
| Regulatory MOU deadline | 60 days (mutual assistance / emergency roaming) |
Magnitude profile
National telecom/IP core collapse: mobile, internet, 911, and Interac payments down ~15h for ~12M subscribers — sub-scores ESTIMATED from public impact reporting, pending deep research.
Sequence of events (SOE)
- TRIGGER During phase six of a seven-phase core IP-network upgrade, Rogers staff remove the Access Control List (ACL) policy filter from the distribution routers' configuration.
- CASCADE With the filter gone, all possible Internet routes pass into the core; a single distribution router releases over 900,000 route entries into the core routers versus ~10,000 normally.
- CASCADE The route flood exceeds the core routers' processing capacity; the core network routers crash within minutes of the filter's removal.
- DETECTION External BGP telemetry (Cloudflare) records a spike in BGP updates after 08:15 UTC, peaking at 08:45 UTC, with a withdrawal of prefixes from the Rogers ASN — Rogers vanishes from the global routing table.
- DETECTION Rogers is internally blind: the outage severs internal access to systems including the VPN to core network nodes, hampering employees' ability to mobilize a team and identify the issue.
- IMPACT Nationwide outage onset — more than 12 million customers lose wireless and wireline services (mobile, home Internet, corporate and institutional).
- IMPACT 9-1-1 access from Rogers mobile phones is unavailable; a Hamilton man reportedly could not call 911 as his sister was dying, though it is unclear whether the death could have been avoided had 911 contact been possible.
- CASCADE Interac is taken offline by the outage, preventing businesses nationwide from accepting debit-card transactions even on other ISPs; some stores temporarily close.
- MITIGATION No suppression analog activates: core routers had no overload protection and the converged architecture had no automatic partition, so nothing contained the flood — a catastrophic loss of all services.
- MITIGATION Recovery cannot be a simple rollback; because the VPN to core nodes is lost, responders must regain management reach, then isolate the flooding distribution routers, purge bad routing state, and rebuild core route tables in stages.
- RECOVERY Cloudflare observes a first attempt to re-advertise Rogers prefixes as the staged rebuild begins.
- RECOVERY 'Another round of withdraws at around 15:45' — Rogers prefixes are withdrawn again, a recovery setback amid core-network flapping.
- RECOVERY Partial recovery of traffic from the Rogers network, mostly after 00:15 UTC, as services are rebuilt in stages.
- RECOVERY Traffic reaches around 76% of the previous day's level at the same time, though frequent BGP announcements/withdrawals show the core flapping issue is not yet fully resolved.
- RESTORED CRTC dates full restoration; the outage window runs 04:58 EDT 8 July to 07:00 EDT 9 July (approximately 26 hours).
- RESTORED Minister Champagne sets a 60-day deadline for carriers to agree mutual assistance and emergency roaming; Rogers pledges $261 million to physically separate its wireless and wireline networks.
Root cause
Contributing factors
- Maintenance/change-control lapse: the ACL policy filter was removed from distribution-router configuration during phase six of a seven-phase core upgrade, with no safeguard against the resulting route flood (CRTC/Xona; Rogers' CRTC letter via Wikipedia).
- Inadequate pre-deployment testing: the change reached production without adequate staged lab validation; the CRTC recommended Rogers 'ensure wide coverage and rigor in testing configuration changes' (CRTC/Xona).
- No overload protection on core routers: 'the Rogers core network routers were not configured with such overload protection mechanisms,' so >900,000 routes (vs ~10,000 normal) crashed the core within minutes (CRTC/Xona).
- Converged single-core architecture: wireless and wireline shared one IP core, so the outage 'resulted in a catastrophic loss of all services' — one change felled mobile, wireline, and 9-1-1 together (CRTC/Xona).
- No independent management path: Rogers lacked redundant NOC/site connectivity from alternative providers, and the outage severed internal access including 'the company's VPN to its core network nodes,' delaying responders (CRTC/Xona; Rogers' CRTC letter via Wikipedia).
- No inter-carrier emergency-roaming failover: Rogers mobile users had no 9-1-1 fallback during the outage, prompting the post-incident emergency-roaming mandate (CRTC/Xona; Champagne undertakings via Wikipedia).
Correction of errors (COE)
- Physically separate wireless and wireline core networks ($261M pledge) so one fault cannot collapse all services
- Implement router overload protection (max-prefix / route limits / control-plane policing) on core routers
- Provision independent/out-of-band management and redundant NOC connectivity from alternative providers
- Strengthen configuration-change control with wide-coverage, rigorous pre-deployment lab testing and staged validation
- Agree inter-carrier mutual-assistance and emergency-roaming arrangements within 60 days
- Establish and regularly test 9-1-1 emergency-roaming failover across carriers
Lessons learnt
- A single removed guardrail in a converged core is a national single point of failure: removing one ACL filter injected over 900,000 routes and crashed the core within minutes, taking mobile, wireline, 9-1-1, and Interac down together (CRTC/Xona).
- The management plane must survive the data plane it manages: because the VPN to core nodes rode the failed network, Rogers was blinded and could not quickly mobilize or roll back — independent, out-of-band access is not optional (CRTC/Xona; Rogers' CRTC letter).
- Overload protection must be assumed-on, not assumed-safe: core routers with no route-limit safeguard forwarded themselves to death when a foreseeable misconfiguration occurred (CRTC/Xona).
- Change control needs rigorous staged lab validation and tested rollback: pushing an unfiltered configuration into production during phase six of a live seven-phase upgrade converted a routine change into a ~26-hour national outage (CRTC/Xona; Rogers' CRTC letter).
- Critical-infrastructure dependencies are systemic: Interac and 9-1-1 riding on one carrier pushed the failure into nationwide debit payments and emergency calling, well beyond Rogers' own subscribers (Wikipedia).
- Vendor detail can be a documented gap: the failed routers' make, model, firmware, and age were never publicly disclosed, so this analysis records them as unknown rather than inventing them (CRTC/Xona; public filings).
Improvements & remediation
- DesignPhysically separate the wireless and wireline core networks so a single fault cannot collapse all services simultaneously — Rogers pledged $261 million toward this separation (Wikipedia; CRTC/Xona rationale).
- DesignImplement router overload-protection mechanisms (max-prefix / route limits / control-plane policing) on core routers so a route flood is rejected rather than allowed to crash the control plane (CRTC/Xona).
- DesignProvide a management path physically independent of the data network, plus redundant NOC/site connectivity from alternative service providers, so responders retain access even when the core fails (CRTC/Xona).
- MaintenanceStrengthen configuration-change control with rigorous, wide-coverage pre-deployment lab testing and staged validation of every core-router change, with tested rollback before production rollout (CRTC/Xona).
- ProcessEstablish an inter-carrier framework for mutual assistance, emergency roaming for rivals' affected customers, and coordinated public communication, on the 60-day deadline set by Minister Champagne (Wikipedia; CRTC/Xona).
- SafetyGuarantee 9-1-1 continuity via inter-operator emergency-roaming failover, tested regularly, so mobile emergency calling survives a single carrier's core outage (CRTC/Xona; Wikipedia).
Comprehensive analysis
No fire — a congestive control-plane collapse
This was not a physical or thermal event; it was a routing/control-plane failure. The 'ignition' analog is the deletion of the ACL policy filter on the distribution routers, and the 'fuel' is the uncontrolled injection of routing state into the core. With the filter present, roughly 10,000 routes reached the core; with it gone, a single distribution router pushed 'over 900,000 route data into the core routers,' exceeding processing capacity and crashing the core 'within minutes.' The mechanism is specific and vendor-agnostic: route-table saturation of the control plane, not any hardware defect.
Latent causes: convergence, no overload protection, no independent management
Three latent conditions turned a foreseeable misconfiguration into a national outage. Convergence put wireless and wireline on one IP core, so the CRTC found the fault 'resulted in a catastrophic loss of all services' with no partition to contain it. Core routers had no overload protection ('were not configured with such overload protection mechanisms'), so nothing rejected the flood. And the management path was not independent — Rogers lacked redundant NOC connectivity and lost 'the company's VPN to its core network nodes,' blinding responders and preventing a quick rollback.
External signature and staged recovery
Cloudflare's BGP telemetry provides an independent, timestamped view: a spike in BGP updates 'after 08:15, reaching its peak at 08:45,' then 'a withdrawal of prefixes from Rogers ASN' (UTC) — Rogers vanishing from the global table. Recovery was halting: a re-advertisement attempt at 14:30 UTC, 'another round of withdraws at around 15:45,' partial traffic 'mostly after 00:15 UTC,' and about 76% of prior-day traffic by 08:40 UTC on 9 July, with continued flapping. This matches the CRTC's ~26-hour window (04:58 EDT 8 Jul to 07:00 EDT 9 Jul) and the need to isolate routers, purge state, and rebuild in stages.
Systemic impact and regulatory response
Because critical services rode Rogers' single core, the blast radius exceeded its own subscribers: more than 12 million customers lost service, Interac debit went offline 'nationwide,' and 9-1-1 was unavailable from Rogers mobile — with a reported Hamilton death whose avoidability is unclear. The response was structural: Rogers pledged $261 million to physically separate wireless and wireline, and Minister Champagne gave carriers 60 days to agree mutual assistance and emergency roaming, with the CRTC/Xona assessment recommending overload protection, independent management, and rigorous change testing.
Disclosed gaps
The router vendor, model, firmware version, and hardware age are not established by any public source and are recorded as unknown rather than inferred. The exact wall-clock time of the ACL removal is not published; it is bounded only as shortly before the 04:58 EDT onset and consistent with Cloudflare's ~08:15-08:45 UTC BGP activity. The 'substantial restoration within ~15-19 hours' figure from the source draft is unsupported and was removed; the defensible measure is the ~26-hour outage window and the 76%-by-08:40-UTC recovery point.
Technical deep-dive
References & provenance
- regulatory CRTC — Independent (Xona) assessment of the July 2022 Rogers outage“When this policy filter was removed, a single distribution router released over 900,000 route data into the core routers.”https://crtc.gc.ca/eng/publications/reports/xona2024.htm
- vendor-status Cloudflare — Cloudflare's view of the Rogers Communications outage in Canada“there was a clear spike in BGP (Border Gateway Protocol) updates after 08:15, reaching its peak at 08:45 ... at 08:45 there was a withdrawal of prefixes from Rogers ASN.”https://blog.cloudflare.com/cloudflares-view-of-the-rogers-communications-outage-in-canada/
- press Wikipedia — 2022 Rogers Communications outage (aggregating Rogers' CRTC letter, CBC, government response)“the deletion of a routing filter on its distribution routers caused all possible routes to the internet to pass through the routers, exceeding the capacity of the routers on its core network ... during the sixth phase in a seven-phase update.”https://en.wikipedia.org/wiki/2022_Rogers_Communications_outage
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).