← All incidents
Incident dossier · Rank #13

Rogers Communications Canada Nationwide Network Outage (July 2022)

Rogers Communications 2022-07-08 15h 0m core impact NetworkSoftwareHuman error

A maintenance change to Rogers' core IP network removed a routing filter, flooding the routers and collapsing the entire national network — taking mobile, internet, 911 emergency calling and Interac payments offline across Canada for roughly 15 hours and affecting about 12 million subscribers.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2022-07-08)Trigger · Network2022-07-082022-07-08Primary fault at Rogers Communications — Rogers national IP core networkRogersRogers national IP core networkRogers national IP core networkDownstream service degraded by the fault: Mobile voice/dataMobile voice/dataDownstream service degraded by the fault: Home internetHome internetDownstream service degraded by the fault: 911 emergency calling911 emergency callingDownstream service degraded by the fault: Interac debit/e-transfer paymentsInterac debit/e-transferpayments…

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Rogers Communications
Data center
Rogers national IP core network
Location
Toronto, Canada
Date
2022-07-08

Impact & scale

Users affected
~12 million subscribers nationwide; 911 and Interac payment networks down
Financial
Regulator-ordered credits + reputational cost; exact figure not consolidated
Scope
Sev-1 national network core (multi-hour)
Services / systems down
  • Mobile voice/data
  • Home internet
  • 911 emergency calling
  • Interac debit/e-transfer payments

Impact data & metrics

Route flood magnitude (one distribution router)Over 900,000 routes injected vs ~10,000 normally
Customers affectedMore than 12 million (wireless + wireline)
Core router crash time after filter removalWithin minutes
Outage window (onset to full restoration)04:58 EDT 8 Jul → 07:00 EDT 9 Jul (~26 hours)
BGP update spike / peak (external detection)Spike after 08:15 UTC, peak 08:45 UTC; prefix withdrawal from Rogers ASN
Traffic recovery by 09 Jul 08:40 UTC~76% of prior-day level at same time
Upgrade phase at failurePhase 6 of a 7-phase core IP-network update
Reported deaths (911 unavailable)At least one reported (avoidability unclear)
Wireless/wireline core separation investment$261 million (CAD) pledged
Regulatory MOU deadline60 days (mutual assistance / emergency roaming)

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 9Users affected (0–10) — breadth of the user/customer population impacted. — scored 9/10.Financial 7Financial impact (0–10) — direct + consequential cost. — scored 7/10.Duration 7Outage duration (0–10) — how long service was degraded/down. — scored 7/10.Blast 9Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 9/10.
Magnitude 8.2 = blast 9×0.35 + users 9×0.25 + financial 7×0.20 + duration 7×0.20 (sub-scores 0–10 · weighted composite)

National telecom/IP core collapse: mobile, internet, 911, and Interac payments down ~15h for ~12M subscribers — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events2022-07-08 pre-onset (approx) · TRIGGER — During phase six of a seven-phase core IP-network upgrade, Rogers staff remove the Access Control List (ACL) policy filter from the distribution routers' configuration.TRIGGER2022-07-08 pre-oimmediately after removal · CASCADE — With the filter gone, all possible Internet routes pass into the core; a single distribution router releases over 900,000 route entries into the core routers versus ~10,000 normally.CASCADEimmediately aftewithin minutes of removal · CASCADE — The route flood exceeds the core routers' processing capacity; the core network routers crash within minutes of the filter's removal.CASCADEwithin minutes o08:15-08:45 UTC (04:15-04:45 EDT) · DETECTION — External BGP telemetry (Cloudflare) records a spike in BGP updates after 08:15 UTC, peaking at 08:45 UTC, with a withdrawal of prefixes from the Rogers ASN — Rogers vanishes from the global routing table.DETECTION08:15-08:45 Uduring onset · DETECTION — Rogers is internally blind: the outage severs internal access to systems including the VPN to core network nodes, hampering employees' ability to mobilize a team and identify the issue.DETECTIONduring onset04:58 EDT · IMPACT — Nationwide outage onset — more than 12 million customers lose wireless and wireline services (mobile, home Internet, corporate and institutional).IMPACT04:58 ED08 Jul, morning · IMPACT — 9-1-1 access from Rogers mobile phones is unavailable; a Hamilton man reportedly could not call 911 as his sister was dying, though it is unclear whether the death could have been avoided had 911 contact been possible.IMPACT08 Jul, morning08 Jul, daytime · CASCADE — Interac is taken offline by the outage, preventing businesses nationwide from accepting debit-card transactions even on other ISPs; some stores temporarily close.CASCADE08 Jul, daytime08 Jul · MITIGATION — No suppression analog activates: core routers had no overload protection and the converged architecture had no automatic partition, so nothing contained the flood — a catastrophic loss of all services.MITIGATION08 Jul08 Jul · MITIGATION — Recovery cannot be a simple rollback; because the VPN to core nodes is lost, responders must regain management reach, then isolate the flooding distribution routers, purge bad routing state, and rebuild core route tables in stages.MITIGATION08 Jul08 Jul 14:30 UTC · RECOVERY — Cloudflare observes a first attempt to re-advertise Rogers prefixes as the staged rebuild begins.RECOVERY08 Jul 14:30 U08 Jul ~15:45 UTC · RECOVERY — 'Another round of withdraws at around 15:45' — Rogers prefixes are withdrawn again, a recovery setback amid core-network flapping.RECOVERY08 Jul ~15:45 U09 Jul after 00:15 UTC · RECOVERY — Partial recovery of traffic from the Rogers network, mostly after 00:15 UTC, as services are rebuilt in stages.RECOVERY09 Jul after 00:15 U09 Jul 08:40 UTC · RECOVERY — Traffic reaches around 76% of the previous day's level at the same time, though frequent BGP announcements/withdrawals show the core flapping issue is not yet fully resolved.RECOVERY09 Jul 08:40 U09 Jul 07:00 EDT · RESTORED — CRTC dates full restoration; the outage window runs 04:58 EDT 8 July to 07:00 EDT 9 July (approximately 26 hours).RESTORED09 Jul 07:00 EDPost-incident (2022) · RESTORED — Minister Champagne sets a 60-day deadline for carriers to agree mutual assistance and emergency roaming; Rogers pledges $261 million to physically separate its wireless and wireline networks.RESTOREDPost-incident (2

Root cause

The direct trigger was a change-management action, not a physical ignition: during the sixth phase of a seven-phase upgrade to Rogers' core IP network, staff removed the Access Control List (ACL) policy filter from the configuration of the distribution routers. The CRTC-commissioned independent (Xona) assessment states "Rogers staff removed the Access Control List policy filter from the configuration of the distribution routers," and quantifies the result: about 10,000 routes are advertised into the core with the filter present, but "when this policy filter was removed, a single distribution router released over 900,000 route data into the core routers." That flood exceeded processing capacity and "the core network routers crashed within minutes from the time the policy filter was removed" — a control-plane/route-table congestive collapse. No fire, combustion, or chemistry was involved; the "ignition source" analog is the ACL deletion and the "fuel" is the uncontrolled route-table injection that saturated the core routers. The specific equipment implicated was the Rogers IP-core routing tier — distribution routers feeding the core network routers. Critically, the router vendor, model, firmware version, and hardware age are NOT publicly disclosed; they were not identified in the public CRTC/Xona assessment or Rogers' CRTC filings, so make/model/age cannot be traced to a source and are recorded here as unknown. The latent root causes are design, redundancy, and maintenance/testing lapses that turned a single foreseeable misconfiguration into a national collapse. Design: the converged architecture ran wireless and wireline over one shared IP core, so per the CRTC "the scope of the outage was extreme in that it resulted in a catastrophic loss of all services." Overload protection: "the Rogers core network routers were not configured with such overload protection mechanisms," so nothing rejected the flood. Management redundancy: Rogers "did not provision its network operation centre and other critical remote infrastructure sites with redundant connectivity from alternative service providers," and the outage severed internal access including "the company's VPN to its core network nodes," blinding responders. Maintenance/testing: the change was pushed to production without adequate staged lab validation — the CRTC recommended Rogers "ensure wide coverage and rigor in testing configuration changes" — with no rollback containment, fault isolation, or inter-operator emergency-roaming failover for 9-1-1.

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

No fire — a congestive control-plane collapse

This was not a physical or thermal event; it was a routing/control-plane failure. The 'ignition' analog is the deletion of the ACL policy filter on the distribution routers, and the 'fuel' is the uncontrolled injection of routing state into the core. With the filter present, roughly 10,000 routes reached the core; with it gone, a single distribution router pushed 'over 900,000 route data into the core routers,' exceeding processing capacity and crashing the core 'within minutes.' The mechanism is specific and vendor-agnostic: route-table saturation of the control plane, not any hardware defect.

Latent causes: convergence, no overload protection, no independent management

Three latent conditions turned a foreseeable misconfiguration into a national outage. Convergence put wireless and wireline on one IP core, so the CRTC found the fault 'resulted in a catastrophic loss of all services' with no partition to contain it. Core routers had no overload protection ('were not configured with such overload protection mechanisms'), so nothing rejected the flood. And the management path was not independent — Rogers lacked redundant NOC connectivity and lost 'the company's VPN to its core network nodes,' blinding responders and preventing a quick rollback.

External signature and staged recovery

Cloudflare's BGP telemetry provides an independent, timestamped view: a spike in BGP updates 'after 08:15, reaching its peak at 08:45,' then 'a withdrawal of prefixes from Rogers ASN' (UTC) — Rogers vanishing from the global table. Recovery was halting: a re-advertisement attempt at 14:30 UTC, 'another round of withdraws at around 15:45,' partial traffic 'mostly after 00:15 UTC,' and about 76% of prior-day traffic by 08:40 UTC on 9 July, with continued flapping. This matches the CRTC's ~26-hour window (04:58 EDT 8 Jul to 07:00 EDT 9 Jul) and the need to isolate routers, purge state, and rebuild in stages.

Systemic impact and regulatory response

Because critical services rode Rogers' single core, the blast radius exceeded its own subscribers: more than 12 million customers lost service, Interac debit went offline 'nationwide,' and 9-1-1 was unavailable from Rogers mobile — with a reported Hamilton death whose avoidability is unclear. The response was structural: Rogers pledged $261 million to physically separate wireless and wireline, and Minister Champagne gave carriers 60 days to agree mutual assistance and emergency roaming, with the CRTC/Xona assessment recommending overload protection, independent management, and rigorous change testing.

Disclosed gaps

The router vendor, model, firmware version, and hardware age are not established by any public source and are recorded as unknown rather than inferred. The exact wall-clock time of the ACL removal is not published; it is bounded only as shortly before the 04:58 EDT onset and consistent with Cloudflare's ~08:15-08:45 UTC BGP activity. The 'substantial restoration within ~15-19 hours' figure from the source draft is unsupported and was removed; the defensible measure is the ~26-hour outage window and the 76%-by-08:40-UTC recovery point.

Technical deep-dive

The failure was a route-table congestive collapse propagating through a converged IP core. In normal operation Rogers' distribution routers passed a filtered, bounded set of routes into the core — the CRTC/Xona assessment states "about 10,000 routes are advertised into the core router when the Access Control List policy filter is present on the distribution router" — with the ACL filter acting as the guardrail that constrained what routing information could enter the core. Phase six of a seven-phase core upgrade removed that ACL filter from the distribution-router configuration. Rogers told the CRTC that "the deletion of a routing filter on its distribution routers caused all possible routes to the internet to pass through the routers, exceeding the capacity of the routers on its core network." The assessment quantifies one distribution router alone injecting "over 900,000 route data into the core routers," and "the core network routers crashed within minutes from the time the policy filter was removed." Because the crashed core carried Rogers' external reachability, the ASN's global advertisements collapsed with it. Cloudflare's external BGP telemetry captured the signature from outside the network: "there was a clear spike in BGP (Border Gateway Protocol) updates after 08:15, reaching its peak at 08:45," and "at 08:45 there was a withdrawal of prefixes from Rogers ASN" (times UTC; 08:45 UTC = 04:45 EDT). That prefix withdrawal is why the outage was total and instantaneous from the public Internet's perspective — Rogers effectively disappeared from the global routing table. Cloudflare assessed the pattern as "likely to be an internal error, not a cyber attack." Two architectural facts made a control-plane bug catastrophic rather than contained. First, convergence: wireless and wireline shared one IP core, so there was no partition to absorb the fault — mobile, home Internet, enterprise, 9-1-1, and Interac dropped simultaneously ("a catastrophic loss of all services"). Second, no independent management path: once the core crashed, Rogers employees lost "internal access to systems such as the company's VPN to its core network nodes, hampering the ability of the company's employees to mobilize a team and identify the issue." Recovery therefore could not be a simple rollback; responders first had to regain management reach, isolate the flooding distribution routers, purge the injected routing state, and rebuild the core route tables and services in stages. Cloudflare's telemetry shows the halting recovery: an advertisement attempt "at 14:30," "another round of withdraws at around 15:45" (a setback), "partial recovery of traffic from the Rogers network, mostly after 00:15 UTC" on 9 July, and "around 76% of the previous day traffic at the same time" by 08:40 UTC on 9 July — with Cloudflare still seeing "frequent BGP announcement and withdrawals" (core "flapping" not fully resolved), consistent with the CRTC dating full restoration to 07:00 EDT on 9 July.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home