← All incidents
Incident dossier · Rank #23

Optus Australia Nationwide Outage — Routing Change Cascade (November 8, 2023)

Optus (Singtel) 2023-11-08 14h 0m core impact NetworkSoftware

Following a routine Singtel-network software upgrade, a large, unexpected flood of routing information exceeded preset safety limits on Optus routers, which self-isolated and cascaded into a roughly 14-hour nationwide outage of Optus mobile and internet for about 10 million Australians — disrupting triple-zero (000) emergency calls, hospitals, transport and payments.

Failure cascade

Failure cascade: trigger → fault → downstream impactTriggerPrimary faultDownstream impactTrigger — Network (2023-11-08)Trigger · Network2023-11-082023-11-08Primary fault at Optus (Singtel) — Optus national network coreOptusOptus national network coreOptus national network coreDownstream service degraded by the fault: Optus mobile voice/data nationwideOptus mobile voice/datanationwide…Downstream service degraded by the fault: Home internetHome internetDownstream service degraded by the fault: Triple-zero (000) emergency calling (partial)Triple-zeroDownstream service degraded by the fault: Transport/payment dependenciesTransport/payment dependencies

Trigger → primary fault → downstream blast radius, derived from the sourced root cause and affected-services record.

Facility & location

Operator
Optus (Singtel)
Data center
Optus national network core
Location
Sydney, Australia
Date
2023-11-08

Impact & scale

Users affected
~10 million Optus customers nationwide; triple-zero (000) calls affected
Financial
Regulatory scrutiny; CEO resignation; customer compensation
Scope
Sev-1 national network core (routing)
Services / systems down
  • Optus mobile voice/data nationwide
  • Home internet
  • Triple-zero (000) emergency calling (partial)
  • Transport/payment dependencies

Impact data & metrics

BGP route-announcement flood~940,000 announcements in one hour vs <3,000/hr normal (~300x)
Provider Edge routers self-disconnected~90 edge/PE routers
People affected>10 million people and 400,000 businesses
Outage duration~12-13 hours (some media reported up to 14)
Restoration progress~11% by 11:30 AEDT; ~88% by 13:00 AEDT
Onset time8 Nov 2023, ~04:05 AEDT
Singtel market-value loss>AU$2 billion (4.8% share-price drop)
Physical-world cascade~500 Melbourne train services cancelled; network shutdown 05:00-06:00 AEDT
Customer compensation200 GB extra data (post-paid); unlimited weekend data (prepaid, rest of 2023)
Emergency-services impactTriple-Zero (000) calls failed (exact count not in retrievable record)
Competitor diversion (press, unverified)TPG/Vodafone ~4x activity; Kogan Mobile +400% eSIM sales
Optus market position (press, unverified)~10 million customers, ~31% market share

Magnitude profile

Magnitude sub-scores (0–10)Magnitude sub-scores (0–10)Users 8Users affected (0–10) — breadth of the user/customer population impacted. — scored 8/10.Financial 6Financial impact (0–10) — direct + consequential cost. — scored 6/10.Duration 6Outage duration (0–10) — how long service was degraded/down. — scored 6/10.Blast 8Blast radius (0–10) — how wide the fault propagated across systems/regions. — scored 8/10.
Magnitude 7.2 = blast 8×0.35 + users 8×0.25 + financial 6×0.20 + duration 6×0.20 (sub-scores 0–10 · weighted composite)

A routing change after a network-software upgrade cascaded into a ~14h nationwide outage for ~10M Australians, disrupting triple-zero (000) — sub-scores ESTIMATED from public impact reporting, pending deep research.

Sequence of events (SOE)

Phased sequence of events7-8 Nov 2023 (pre-onset) · TRIGGER — Scheduled software upgrade performed on a router at a North American node of Singtel's international peering network, the upstream infrastructure Optus's IP core depends on.TRIGGER7-8 Nov 2023 (pr8 Nov 2023, ~04:05 AEDT · TRIGGER — The Singtel exchange upgrade caused one router to disconnect, initiating an abnormal BGP route-advertisement event toward Optus's AS4804.TRIGGER8 Nov 2023, ~04:05 AED8 Nov 2023, ~04:05 AEDT · DETECTION — Cloudflare telemetry for AS4804 registered the anomaly as a mass BGP spike — 'over 940,000 announcements in an hour from a node that normally makes less than 3,000 announcements per hour.'DETECTION8 Nov 2023, ~04:05 AED8 Nov 2023, ~04:05 AEDT · DETECTION — Optus's Cisco PE routers registered the flood as incoming route updates crossed 'the pre-configured default threshold limits set by Cisco Systems.'DETECTION8 Nov 2023, ~04:05 AED8 Nov 2023, ~04:05 AEDT · MITIGATION — Built-in threshold self-protection activated as designed: routers began tearing down BGP sessions to prevent CPU/memory exhaustion.MITIGATION8 Nov 2023, ~04:05 AED8 Nov 2023, ~04:05-04:10 AEDT · CASCADE — 'Approximately 90 edge provider routers disconnected as an automated protective measure against routing update overload' — a synchronized self-isolation across the PE fleet.CASCADE8 Nov 2023, ~04:05-04:10 AED8 Nov 2023, ~04:10 AEDT · CASCADE — With ~90 PE routers withdrawn from BGP, Optus's IP core effectively vanished from the internet for its customers (network-layer 'de-energisation').CASCADE8 Nov 2023, ~04:10 AED8 Nov 2023, morning · IMPACT — Nationwide loss of mobile and fixed service affecting 'more than 10 million people' and '400,000 businesses.'IMPACT8 Nov 2023, morn8 Nov 2023, morning · IMPACT — Emergency Triple-Zero (000) calls failed for affected customers; exact count of failed calls not in retrievable record.IMPACT8 Nov 2023, morn8 Nov 2023, 05:00-06:00 AEDT · IMPACT — Physical-world cascade: Melbourne's train network experienced a shutdown from 05:00 to 06:00 and about 500 train services were cancelled; hospital phone lines and EFTPOS/payment terminals were disrupted.IMPACT8 Nov 2023, 05:00-06:00 AED8 Nov 2023, daytime · CASCADE — Customers diverted to rivals: TPG/Vodafone reported a roughly four-fold activity increase and Kogan Mobile a ~400% rise in eSIM sales (press figures, not re-verified this session).CASCADE8 Nov 2023, dayt8 Nov 2023, daytime · MITIGATION — Optus publicly confirmed the outage was 'not due to a cyberattack,' focusing recovery on the routing fault.MITIGATION8 Nov 2023, dayt8 Nov 2023, morning-afternoon · RECOVERY — Because tripped BGP sessions do not self-heal, engineers manually reconnected/rebooted affected routers (manual/on-site detail inferred from the recovery profile).RECOVERY8 Nov 2023, morn8 Nov 2023, ~11:30 AEDT · RECOVERY — Partial restoration reached approximately 11% of the network.RECOVERY8 Nov 2023, ~11:30 AED8 Nov 2023, ~13:00 AEDT · RECOVERY — Restoration advanced to approximately 88% of the network.RECOVERY8 Nov 2023, ~13:00 AED8 Nov 2023, ~evening · RESTORED — Full service restored after approximately 12-13 hours (some media reported up to a 14-hour outage).RESTORED8 Nov 2023, ~eveNov 2023 (aftermath) · IMPACT — Singtel lost over AU$2 billion in value, a 4.8% share-price drop attributed to the outage.IMPACTNov 2023 (afterm20 Nov 2023 · IMPACT — Optus CEO Kelly Bayer Rosmarin resigned; a Senate public inquiry and an independent government (Bean) review were established.IMPACT20 Nov 2023

Root cause

The proximate ignition was a scheduled software upgrade on a router at a North American node of Singtel's international peering network — the upstream infrastructure Optus's IP core depends on as a Singtel subsidiary. Per the retrievable record, "a software upgrade at a North American Singtel exchange caused one router to disconnect," and that disconnection acted as the spark: it emitted an abnormal flood of BGP routing advertisements toward Optus's autonomous system AS4804. Cloudflare telemetry (as cited by Wikipedia) recorded "over 940,000 announcements in an hour from a node that normally makes less than 3,000 announcements per hour" — roughly a 300-fold surge. (Note: some enriched sourcing named this the "STiX/Singtel Internet Exchange" fabric; the retrievable record confirms only "a North American Singtel exchange," so the specific fabric name is not established here.) The specific vulnerable equipment was Optus's fleet of Provider Edge (PE) routers, confirmed to be Cisco Systems devices. The failure mechanism was a BGP route-limit self-protection trip: the flood caused Optus routers to "rapidly update routing tables and exceed the pre-configured default threshold limits set by Cisco Systems," and each router responded by executing the standard protective action of tearing down its BGP session to avoid CPU/memory exhaustion. "Approximately 90 edge provider routers disconnected as an automated protective measure against routing update overload." The safety mechanism performed exactly as engineered — but its activation IS the outage: by isolating themselves, the ~90 PE routers severed Optus customers from the network. The exact Cisco model, firmware and hardware age are not disclosed in the retrievable record. The latent root causes are two change-management/configuration design lapses evidenced by the technical description. First, an upstream peering-hygiene gap: a Singtel exchange upgrade was permitted to leak a ~940,000-route flood downstream into Optus with no apparent prefix-filtering or rate-limiting to contain it — a foreseeable containment gap at the peering boundary (inferred from the mechanism, not a verbatim finding). Second, an Optus-side configuration lapse: its PE routers were left on Cisco's DEFAULT threshold ceilings rather than thresholds tuned to Optus's real route counts and headroom, so a foreseeable flood tripped a hard fail-safe (session teardown) rather than degrading gracefully. Because the identical default limit tripped near-simultaneously across ~90 PE routers, the "containment" produced a synchronized disconnection cascade rather than a firebreak. Attribution of change ownership was reportedly disputed (Singtel was reported to have refuted early causation claims; not re-verified this session). Optus expressly ruled out a cyberattack. The independent government (Bean) review and the Senate/ACMA processes were established to examine these change-management, Triple-Zero failover and recovery shortcomings; their verbatim findings were NOT directly retrievable this session (government servers returned HTTP 403 / timed out).

Contributing factors

Correction of errors (COE)

Lessons learnt

Improvements & remediation

Comprehensive analysis

What actually failed

A scheduled software upgrade on a router at a North American Singtel exchange caused that router to disconnect, generating an abnormal BGP re-advertisement event toward Optus's AS4804 — Cloudflare-cited telemetry recorded over 940,000 announcements in one hour versus a normal baseline below 3,000. Optus's Cisco Provider Edge routers hit 'the pre-configured default threshold limits set by Cisco Systems' and, as designed, tore down their BGP sessions; approximately 90 edge routers disconnected, effectively removing Optus's IP core from the internet. The protective mechanism working correctly was the outage.

Why a protective feature became the disaster

Two latent design lapses turned a foreseeable event into a national outage. The PE fleet ran on Cisco DEFAULT thresholds rather than limits tuned to Optus's real route counts, so a ~300x flood breached a hard fail-safe instead of degrading gracefully. Because the identical default limit tripped across ~90 routers at once, the 'containment' produced a synchronized cascade with no firebreak. Upstream, the Singtel exchange upgrade propagated the route event downstream with no evident prefix-filtering or rate-limiting at the peering boundary. Both are inferences from the confirmed mechanism, not verbatim official findings.

Impact and physical-world cascade

Service loss hit more than 10 million people and 400,000 businesses. Triple-Zero (000) emergency calls failed — the issue that dominated the regulatory response. Melbourne's train network shut down between 05:00 and 06:00 with about 500 services cancelled; hospital phone lines and EFTPOS terminals were disrupted. Restoration was gradual (11% by 11:30, 88% by 13:00 AEDT) over roughly 12-13 hours, consistent with BGP sessions that do not self-heal and require manual re-enablement. Singtel lost over AU$2 billion in value (4.8% share drop) and CEO Kelly Bayer Rosmarin resigned on 20 November 2023.

Evidence quality and open gaps

Only Wikipedia was directly re-fetched this session; the Cloudflare blog (404), ABC, ZDNet, web.archive and Australian government servers (Bean review / ACMA / Senate report) were blocked or timed out. So the Cisco-default-threshold mechanism, ~90 routers, 940,000/3,000 figures, timings, scale, market loss and CEO resignation are confirmed via Wikipedia; competitor-diversion and market-share figures (Motley Fool) and the Singtel causation refutation (ZDNet) remain single-sourced and unverified this session. The verbatim official findings and the exact number of failed 000 calls and Cisco router model/firmware were not retrievable. officialPostmortem is therefore set false: no official/regulatory document was directly obtained with a quotable passage.

Technical deep-dive

This was a routing-layer cascade, not a physical fire, so the forensics map onto network primitives, and the deepest technical layer is presented as informed analysis grounded in the one confirmed fact — Cisco "pre-configured default threshold limits" — rather than as verbatim sourcing. IGNITION: a Singtel exchange software upgrade in North America caused a router to disconnect, which in BGP terms generates a mass re-advertisement/withdrawal event. Cloudflare (via Wikipedia) observed AS4804 emitting "over 940,000 announcements in an hour" against a baseline "less than 3,000 announcements per hour." PROPAGATION (spread): BGP re-advertises reachability hop-by-hop, so the flood propagated automatically and near-instantly across Optus's IP core, with no effective firebreak. SUPPRESSION (self-protection): each Cisco PE router enforces a configured maximum-received-prefix ceiling; when incoming prefixes exceed it, the standard, RFC-aligned protective response is to tear down the offending BGP session to prevent CPU/RIB/FIB exhaustion. Here the ceilings were left at or near Cisco's DEFAULT value, so the flood breached them fleet-wide and "approximately 90 edge provider routers disconnected." DE-ENERGISATION analogue: the ~90 PE routers withdrew their sessions and routes, so Optus's IP core effectively vanished from the internet for its customers — a self-isolation rather than an electrical trip. NON-AUTO-RESTORE: tripped max-prefix sessions typically do not silently re-establish when the flood subsides; standard behaviour holds the session down until an operator re-enables it (often a clear/reset or reboot). This is consistent with restoration requiring manual intervention and taking the better part of a day rather than minutes (the manual, sometimes on-site, reconnection detail is a reasonable inference from the recovery profile, not a verbatim source quote). The mechanism is distinct from the 2022 Rogers outage, where a maintenance change removed a route filter and flooded the core from within; here the trigger originated upstream at Singtel and the failing safeguard was a preset route-count limit firing synchronously across the fleet. Exact router model/firmware, the precise count of failed 000 calls, and welfare-check tallies reside in the Senate report / ACMA / Bean findings, which were NOT directly retrievable this session.

References & provenance

Sourced from public post-incident reports. Quotes are short attributed excerpts for provenance only; the analysis above is original and substantially shorter than its sources. Last verified 2026-08-02 (seed — pending deep research).

Root access required

The DC Incidents dossier is a root-only module. Sign in with an authorized account to continue.

Back to Home