Optus Australia Nationwide Outage — Routing Change Cascade (November 8, 2023)
Following a routine Singtel-network software upgrade, a large, unexpected flood of routing information exceeded preset safety limits on Optus routers, which self-isolated and cascaded into a roughly 14-hour nationwide outage of Optus mobile and internet for about 10 million Australians — disrupting triple-zero (000) emergency calls, hospitals, transport and payments.
Failure cascade
Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.
Facility & location
- Operator
- Optus (Singtel)
- Data center
- Optus national network core
- Location
- Sydney, Australia
- Date
- 2023-11-08
Impact & scale
- Users affected
- ~10 million Optus customers nationwide; triple-zero (000) calls affected
- Financial
- Regulatory scrutiny; CEO resignation; customer compensation
- Scope
- Sev-1 national network core (routing)
- Optus mobile voice/data nationwide
- Home internet
- Triple-zero (000) emergency calling (partial)
- Transport/payment dependencies
Impact data & metrics
| BGP route-announcement flood | ~940,000 announcements in one hour vs <3,000/hr normal (~300x) |
| Provider Edge routers self-disconnected | ~90 edge/PE routers |
| People affected | >10 million people and 400,000 businesses |
| Outage duration | ~12-13 hours (some media reported up to 14) |
| Restoration progress | ~11% by 11:30 AEDT; ~88% by 13:00 AEDT |
| Onset time | 8 Nov 2023, ~04:05 AEDT |
| Singtel market-value loss | >AU$2 billion (4.8% share-price drop) |
| Physical-world cascade | ~500 Melbourne train services cancelled; network shutdown 05:00-06:00 AEDT |
| Customer compensation | 200 GB extra data (post-paid); unlimited weekend data (prepaid, rest of 2023) |
| Emergency-services impact | Triple-Zero (000) calls failed (exact count not in retrievable record) |
| Competitor diversion (press, unverified) | TPG/Vodafone ~4x activity; Kogan Mobile +400% eSIM sales |
| Optus market position (press, unverified) | ~10 million customers, ~31% market share |
Magnitude profile
A routing change after a network-software upgrade cascaded into a ~14h nationwide outage for ~10M Australians, disrupting triple-zero (000) — sub-scores ESTIMATED from public impact reporting, pending deep research.
Sequence of events (SOE)
- TRIGGER Scheduled software upgrade performed on a router at a North American node of Singtel's international peering network, the upstream infrastructure Optus's IP core depends on.
- TRIGGER The Singtel exchange upgrade caused one router to disconnect, initiating an abnormal BGP route-advertisement event toward Optus's AS4804.
- DETECTION Cloudflare telemetry for AS4804 registered the anomaly as a mass BGP spike — 'over 940,000 announcements in an hour from a node that normally makes less than 3,000 announcements per hour.'
- DETECTION Optus's Cisco PE routers registered the flood as incoming route updates crossed 'the pre-configured default threshold limits set by Cisco Systems.'
- MITIGATION Built-in threshold self-protection activated as designed: routers began tearing down BGP sessions to prevent CPU/memory exhaustion.
- CASCADE 'Approximately 90 edge provider routers disconnected as an automated protective measure against routing update overload' — a synchronized self-isolation across the PE fleet.
- CASCADE With ~90 PE routers withdrawn from BGP, Optus's IP core effectively vanished from the internet for its customers (network-layer 'de-energisation').
- IMPACT Nationwide loss of mobile and fixed service affecting 'more than 10 million people' and '400,000 businesses.'
- IMPACT Emergency Triple-Zero (000) calls failed for affected customers; exact count of failed calls not in retrievable record.
- IMPACT Physical-world cascade: Melbourne's train network experienced a shutdown from 05:00 to 06:00 and about 500 train services were cancelled; hospital phone lines and EFTPOS/payment terminals were disrupted.
- CASCADE Customers diverted to rivals: TPG/Vodafone reported a roughly four-fold activity increase and Kogan Mobile a ~400% rise in eSIM sales (press figures, not re-verified this session).
- MITIGATION Optus publicly confirmed the outage was 'not due to a cyberattack,' focusing recovery on the routing fault.
- RECOVERY Because tripped BGP sessions do not self-heal, engineers manually reconnected/rebooted affected routers (manual/on-site detail inferred from the recovery profile).
- RECOVERY Partial restoration reached approximately 11% of the network.
- RECOVERY Restoration advanced to approximately 88% of the network.
- RESTORED Full service restored after approximately 12-13 hours (some media reported up to a 14-hour outage).
- IMPACT Singtel lost over AU$2 billion in value, a 4.8% share-price drop attributed to the outage.
- IMPACT Optus CEO Kelly Bayer Rosmarin resigned; a Senate public inquiry and an independent government (Bean) review were established.
Root cause
Contributing factors
- Configuration/maintenance lapse on Optus PE routers: they were left on 'the pre-configured default threshold limits set by Cisco Systems' rather than thresholds tuned to Optus's real route counts, so a foreseeable flood tripped a hard fail-safe instead of degrading gracefully (Wikipedia, citing the outage account).
- Synchronized fail-safe design flaw: because the identical default limit tripped across ~90 PE routers near-simultaneously, the protective mechanism produced a fleet-wide disconnection cascade rather than isolating a single fault — the 'containment' had no firebreak (Wikipedia).
- Upstream peering-hygiene gap (inferred from mechanism): a North American Singtel exchange upgrade was allowed to propagate a ~940,000-route BGP event downstream toward AS4804 (normal <3,000/hr) with no apparent upstream prefix-filtering or rate-limiting to contain it (Cloudflare figure via Wikipedia).
- No automatic recovery / manual-only restoration: tripped BGP sessions did not self-heal, so restoration proceeded gradually (11% by 11:30, 88% by 13:00) and the outage ran ~12-13 hours (Wikipedia).
- Triple-Zero (000) failover inadequacy: emergency 000 calls failed with no effective resilient fallback during the IP-core loss, later a central regulatory concern of ACMA / the Senate inquiry (Wikipedia).
- Disputed / unclear change ownership at the Singtel-Optus boundary: Singtel was reported to have refuted early claims that its upgrade caused the outage (ZDNet, not re-verified this session), indicating unclear accountability and coordination.
Correction of errors (COE)
- Tune all Provider Edge router route-limit thresholds off Cisco defaults to operator-specific values with graceful-degradation (warn/restart) behaviour instead of synchronized hard teardown
- Implement inbound prefix-filtering and route rate-limiting at the Singtel peering boundary to Optus
- Establish a resilient Triple-Zero (000) failover independent of the Optus IP core plus mandated welfare-check procedures
- Institute joint Singtel-Optus staged change-management with blast-radius review and coordinated rollback
Lessons learnt
- A safety mechanism that fires fleet-wide simultaneously is not containment but a cascade generator: leaving ~90 PE routers on identical Cisco default thresholds turned a protective feature into the outage itself (Wikipedia).
- Default vendor configuration is a latent hazard — thresholds must be tuned to the operator's real traffic, because a foreseeable ~300x route event tripped ceilings never sized for Optus's network (Cloudflare figure via Wikipedia).
- Blast radius crosses corporate boundaries: an upstream Singtel exchange upgrade in North America took down 10M+ Australians, making peering-boundary filtering and joint change-control essential, not optional (Wikipedia).
- Emergency-services resilience must be decoupled from the commercial IP core: 000 call failure, not the commercial outage alone, drove the regulatory and parliamentary response (Wikipedia).
- Recovery design matters as much as failure design: because tripped BGP sessions do not self-heal, a minutes-long trigger became a ~12-13-hour restoration — automated/remote recovery would have sharply cut the impact (Wikipedia).
Improvements & remediation
- DesignReplace Cisco DEFAULT threshold ceilings on all Provider Edge routers with operator-tuned limits sized to real route counts plus headroom, and prefer soft/warning-then-restart behaviour over synchronized hard session-teardown so a flood degrades gracefully instead of self-isolating ~90 routers at once (grounded in the 'pre-configured default threshold limits' failure, Wikipedia).
- MaintenanceEnforce peering-boundary hygiene — inbound prefix-filtering, route-count rate-limiting and AS-path sanity checks at the Singtel ingress — so an upstream exchange upgrade cannot propagate a ~940,000-route event downstream to Optus (Cloudflare figure via Wikipedia).
- Process/Maintenance: Institute joint Singtel-Optus change-management with staged rollout, pre-change blast-radius review and coordinated rollback, closing the disputed accountability gap at the peering boundary (reported Singtel refutation, ZDNet).
- SafetyBuild a resilient Triple-Zero (000) failover path independent of the Optus IP core (e.g. mandated camp-on/roaming to other carriers' networks) plus automated welfare-check procedures, the central concern of ACMA and the Senate inquiry (Wikipedia).
- ProcessAdd automated router-session recovery and pre-staged remote reboot capability so restoration does not depend on gradual manual reconnection that stretched the outage to ~12-13 hours (Wikipedia recovery profile).
Comprehensive analysis
What actually failed
A scheduled software upgrade on a router at a North American Singtel exchange caused that router to disconnect, generating an abnormal BGP re-advertisement event toward Optus's AS4804 — Cloudflare-cited telemetry recorded over 940,000 announcements in one hour versus a normal baseline below 3,000. Optus's Cisco Provider Edge routers hit 'the pre-configured default threshold limits set by Cisco Systems' and, as designed, tore down their BGP sessions; approximately 90 edge routers disconnected, effectively removing Optus's IP core from the internet. The protective mechanism working correctly was the outage.
Why a protective feature became the disaster
Two latent design lapses turned a foreseeable event into a national outage. The PE fleet ran on Cisco DEFAULT thresholds rather than limits tuned to Optus's real route counts, so a ~300x flood breached a hard fail-safe instead of degrading gracefully. Because the identical default limit tripped across ~90 routers at once, the 'containment' produced a synchronized cascade with no firebreak. Upstream, the Singtel exchange upgrade propagated the route event downstream with no evident prefix-filtering or rate-limiting at the peering boundary. Both are inferences from the confirmed mechanism, not verbatim official findings.
Impact and physical-world cascade
Service loss hit more than 10 million people and 400,000 businesses. Triple-Zero (000) emergency calls failed — the issue that dominated the regulatory response. Melbourne's train network shut down between 05:00 and 06:00 with about 500 services cancelled; hospital phone lines and EFTPOS terminals were disrupted. Restoration was gradual (11% by 11:30, 88% by 13:00 AEDT) over roughly 12-13 hours, consistent with BGP sessions that do not self-heal and require manual re-enablement. Singtel lost over AU$2 billion in value (4.8% share drop) and CEO Kelly Bayer Rosmarin resigned on 20 November 2023.
Evidence quality and open gaps
Only Wikipedia was directly re-fetched this session; the Cloudflare blog (404), ABC, ZDNet, web.archive and Australian government servers (Bean review / ACMA / Senate report) were blocked or timed out. So the Cisco-default-threshold mechanism, ~90 routers, 940,000/3,000 figures, timings, scale, market loss and CEO resignation are confirmed via Wikipedia; competitor-diversion and market-share figures (Motley Fool) and the Singtel causation refutation (ZDNet) remain single-sourced and unverified this session. The verbatim official findings and the exact number of failed 000 calls and Cisco router model/firmware were not retrievable. officialPostmortem is therefore set false: no official/regulatory document was directly obtained with a quotable passage.
Technical deep-dive
References & provenance
- press 2023 Optus outage — Wikipedia (root cause and router mechanism)“Approximately 90 edge provider routers disconnected as an automated protective measure against routing update overload ... exceed the pre-configured default threshold limits set by Cisco Systems.”https://en.wikipedia.org/wiki/2023_Optus_outage
- press 2023 Optus outage — overview (Wikipedia, secondary summary of Optus's stated cause)“changes to routing information from an international peering network ... exceeded preset safety levels on key routers, which disconnected from the Optus IP Core network to protect themselves.”https://en.wikipedia.org/wiki/2023_Optus_outage
Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).