🚨 Disaster Communication Network Failure Coordination
This simulation models the coordination of medical assistance during a disaster when communication networks fail, emphasizing effective communication and resource allocation.
Normal Network Topology in EMS Coordination
Modern emergency medical coordination runs almost entirely on infrastructure the public also depends on: commercial cellular networks, fiber-backed internet, and cloud-hosted dispatch software. In steady state this is fast, cheap, and redundant enough for day-to-day 911 calls. It is also, as every major disaster of the last two decades has shown, a single interdependent system with shared failure points.
- ~420,000: US cell sites (approx.) (FCC-reported, 2023)
- 4–8 hrs: Tower backup battery life (typical, without generator)
- 5–8: CAD-to-hospital data hops (dispatch → EHR handoff)
- ~600,000: US daily 911 calls (~240M / year)
What the healthy mesh looks like
In blue-sky operations, every actor in a medical response — hospital emergency departments, ambulance crews, field triage teams, and the Emergency Operations Center (EOC) or Incident Command Post — is reachable through the same commercial pipes: LTE/5G voice and data, computer-aided dispatch (CAD) over the internet, and hospital electronic health record (EHR) systems that ingest ambulance patient data automatically.
This is efficient because it reuses infrastructure nobody has to build or maintain specifically for emergencies. A paramedic's tablet, a hospital's bed-tracking dashboard, and a 911 dispatcher's CAD terminal are all, structurally, just internet endpoints.
The hidden single points of failure
The efficiency comes at a cost: almost every node in the mesh depends on the same handful of physical layers. A cell tower needs commercial grid power (batteries typically bridge only 4–8 hours), a working backhaul link (fiber, microwave, or satellite trunk back to the carrier core), and an intact tower structure. A hospital's internet-based systems depend on the same carrier network, or on wired ISP service running through the same fiber conduits that flood, burn, or snap in a disaster.
Because dispatch, hospital status boards, and field communications all ride the same commercial rails, a single regional outage — a flooded central office, a downed transmission line, a fiber cut from construction equipment — can silently degrade every layer of the response system at once, even though no single piece of medical equipment failed.
FEMA and DHS doctrine formalizes the fix as PACE planning — Primary, Alternate, Contingency, Emergency communication methods — required for every layer of incident communication precisely because commercial networks are treated as a single point of failure, not a guarantee.
Disaster Strikes — Infrastructure Damage & Congestion Collapse
Networks rarely fail because a bomb hits a switch. They fail through an accumulation of ordinary physical and logical stresses arriving simultaneously: lost power, severed fiber, and a public that all reaches for a phone at the same moment. The result is a network that is not destroyed but overwhelmed — a phenomenon disaster communications planners call congestion collapse.
- ~95%: PR cell sites down, Maria (Sept 2017, day 1)
- ~11 months: Full PR restoration (some rural sites)
- 7.0 Mw: Haiti 2010 magnitude (Port-au-Prince)
- 10–50×: Call attempts vs. capacity (typical disaster surge)
Three ways the network actually dies
1. Power loss: cell towers run on grid power with battery or generator backup. Batteries typically last 4–8 hours; generators need fuel resupply that disaster-clogged roads often cannot deliver. When backup runs out, the radio stays physically intact but silent.
2. Physical damage to backhaul: even towers that survive wind or shaking are useless if the fiber or microwave link carrying their traffic back to the carrier core is cut — by falling trees, flooding of underground vaults, or debris. A single severed conduit can silently kill dozens of towers downstream.
3. Congestion collapse: this is the least intuitive failure mode. Surviving capacity is often overwhelmed not by damage but by demand — everyone in an affected area simultaneously tries to call, text, and stream at once. Call volumes can spike 10–50× normal levels within minutes of a disaster, and cellular networks are engineered for average, not panic-driven peak, load. The network appears "down" to users even where towers are fully functional.
Hurricane Maria, Puerto Rico (2017)
Hurricane Maria remains the starkest modern case study. Within a day of landfall, roughly 95% of Puerto Rico's cell sites were out of service — most from a combination of grid power loss (the island's entire power grid collapsed) and physical destruction of towers and fiber. Diesel generators kept some sites alive for days, but fuel logistics collapsed alongside the road network.
Restoration was not measured in hours but in months: while emergency roaming and portable "cell on wheels" (COW) units restored partial service to population centers within weeks, some rural municipalities did not see stable cellular service return for nearly a year. Hospitals lost the ability to transmit patient records, coordinate transfers, or even confirm which facilities still had power and capacity — a purely communications failure that compounded a public health crisis.
FCC data on Hurricane Maria: approximately 95% of Puerto Rico's cell sites were non-operational the day after landfall, and some communities went without functioning cellular service for nearly a year — the longest telecommunications outage in modern US disaster history.
Haiti 2010 — a preview of the same failure
The 2010 Haiti earthquake (Mw 7.0) destroyed or disabled a large share of Port-au-Prince's already-fragile telecommunications infrastructure within seconds. Government communications and most cellular capacity collapsed simultaneously with the healthcare system it needed to coordinate.
What filled the vacuum foreshadowed the entire modern disaster-communications playbook: volunteer amateur ("ham") radio operators — many coordinating through the Haitian amateur radio community and international relief nets — carried emergency traffic when nothing else worked, while a crowdsourced OpenStreetMap effort (built by thousands of remote volunteers from satellite imagery and SMS reports) became the de facto situational map used by responding agencies, because no authoritative digital map of the shattered city existed. Haiti demonstrated that the fallback layer is not theoretical — it is what disaster medicine actually runs on in the first 72 hours.
Fallback to Radio & Satellite Channels
PACE — Primary, Alternate, Contingency, Emergency — is the doctrine that turns communications failure from a crisis into a procedure. When the Primary method (cellular/internet) collapses, trained responders and agencies already know the Alternate: licensed radio and satellite systems that do not depend on commercial grid power or carrier infrastructure at all.
- 66: Iridium satellite count (LEO constellation, global coverage)
- 1–5 mi: VHF/UHF handheld range (line of sight, more via repeater)
- 1950s–: RACES/MARS volunteer nets (FCC/DoD-affiliated emergency radio)
- DHS/CISA: ESF #2 lead agency (National Response Framework)
PACE planning — the doctrine behind the fallback
PACE structures every communications plan into four tiers, each independent of the one before it:
• Primary — the normal method: cellular voice/data, internet-based CAD, EHR systems • Alternate — a different technology reaching the same people: VHF/UHF land mobile radio, ham radio nets • Contingency — a materially different path, often slower: satellite phone, HF (high-frequency) radio, mesh data radios • Emergency — the last resort when nothing electronic works: runners, dispatch riders, physical relay points
The key design principle is technological independence: each tier should not share the failure mode of the tier above it. A satellite phone still works when every cell tower on the ground is dead because it talks directly to orbiting satellites, not to terrestrial infrastructure.
Amateur radio and volunteer emergency nets
VHF/UHF ("2-meter" and "70cm" band) handheld and mobile radios remain one of the most reliable fallback layers because they need no infrastructure beyond a battery and, ideally, a repeater on a hilltop or building to extend range from a few miles to tens of miles.
MARS (Military Auxiliary Radio System) and RACES (Radio Amateur Civil Emergency Service) are formal US programs — dating to the Cold War era — that organize licensed amateur radio operators to provide emergency traffic handling for government agencies when normal systems fail. These are not hobbyists improvising: they operate under FCC Part 97 rules and established emergency-traffic protocols (formal message precedence, NTS — National Traffic System — relay procedures), and are frequently activated during hurricanes, wildfires, and earthquakes to relay hospital status, shelter needs, and casualty counts between EOCs.
ESF #2 (Emergency Support Function #2 — Communications) of the US National Response Framework, coordinated by DHS/CISA, exists specifically to restore and coordinate exactly this layered fallback — deploying portable cell towers, satellite uplinks, and amateur radio liaison teams to disaster zones as a standing federal capability.
Satellite phones and emerging mesh systems
Satellite phone systems like Iridium (66 low-earth-orbit satellites providing true pole-to-pole coverage) let a single handset reach anywhere on Earth with a clear sky view, independent of any terrestrial carrier. Bandwidth is low — enough for voice and short data/text — but availability is nearly unconditional, which is exactly what disaster coordination needs for critical status updates.
Newer mesh technologies are reshaping the Contingency tier: Meshtastic and similar LoRa-based mesh radios let a chain of small, cheap, solar-rechargeable nodes relay short text messages hop-by-hop for miles without any central infrastructure. Starlink and other LEO broadband constellations can restore near-normal internet bandwidth to a hospital or command post within hours of shipping a dish, and were deployed at scale during recent hurricane responses to bridge exactly the gap Hurricane Maria exposed in 2017.
Backup communication methods compared
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Cellular / LTE (intact) | Miles per tower | Mbps — voice, data, video | Grid + 4–8 hr battery; pre-built, none once down |
| VHF/UHF handheld radio | 1–5 mi (10s of mi w/ repeater) | Voice / narrow digital, kbps | Battery, days; minutes to activate |
| HF (ham) radio | 100s–1,000s of mi (skywave) | Voice / digital modes, low kbps | 12V battery or generator; 15–60 min (antenna) |
| Satellite phone (Iridium) | Global, pole-to-pole | 2.4–9.6 kbps voice/data | Battery; minutes, needs sky view |
| Mesh radio / Starlink | Miles (mesh) to global (dish) | Low (LoRa) to high (Mbps) bandwidth | Battery/solar; minutes to hours |
| Human runner | As far as they can travel | Effectively unlimited, extremely slow | None required; immediate but slow |
Runner & Relay Networks — the Oldest Fallback
When radio range, terrain, or damaged repeaters still leave clusters of responders unreachable, the PACE doctrine's final Emergency tier kicks in: a physical human being carries the message. It sounds primitive because it is — and it works precisely because it depends on nothing except a person willing to walk, drive, or bike between two points.
- ~0: Runner failure modes (no batteries, no signal needed)
- 20–90 min: Typical relay message delay (distance-dependent)
- ICS Form 213: ICS message log standard (general message form)
- thousands: Haiti crowdmap contributors (remote OSM volunteers, 2010)
- ancient–present: Applicable historical era (couriers to modern relay points)
Why runners still matter in 21st-century medicine
A human runner carrying a written message packet has no battery to die, no signal to lose, and no frequency to be jammed by congestion. In incidents where isolated clusters of responders or facilities have no working radio link — dead zones behind terrain, a repeater destroyed, or radios simply not distributed that far down the response chain — a runner remains the only channel that is always technically available, if slow.
In formal Incident Command System (ICS) operations, this is not improvisation: message runners are a recognized role, and the standard ICS Form 213 general message form exists specifically to make sure a message physically carried between posts is logged, timestamped, and accountable in the same way a radio transmission would be — sender, recipient, time sent, time received.
Relay points and hub-and-spoke physical networks
Runners rarely travel the full distance between two isolated clusters alone. Response plans typically establish relay points — a staging area, a fixed post, or a vehicle checkpoint — where one runner hands a message packet to another, extending the effective range of the network the way a radio repeater extends radio range, but built entirely out of people and locations.
This physical hub-and-spoke relay was exactly the model that let information move across Port-au-Prince in the days after the 2010 earthquake: messages, patient counts, and supply requests moved between clinics and the coordination hub via a combination of ham radio operators and physical couriers relaying between fixed points, feeding data that was then compiled — often by remote volunteers thousands of miles away — into the crowdsourced OpenStreetMap layer that responding agencies used as their operating map.
The same relay logic scaled to national tragedy at the World Trade Center on September 11, 2001: FDNY radios could not reliably reach commanders inside the North Tower stairwells, and NYPD and FDNY used separate, non-interoperable radio systems entirely. The resulting communications failure is widely cited as a contributing factor in first-responder deaths and became the single biggest driver of the interoperability standards — common channels, unified ICS structure, mandated interagency compatibility — that US emergency communications doctrine uses today.
The tradeoff: throughput versus availability
A runner-based relay has effectively unlimited "bandwidth" in the sense that a written packet can carry pages of detail — patient counts, resource needs, hand-drawn maps — that would take a garbled, congested radio channel far longer to relay verbally. What it sacrifices is latency: a message that would travel at the speed of light over radio instead moves at the speed of a person on foot or in a vehicle over damaged roads, typically adding 20–90 minutes of delay per hop depending on distance and terrain.
This is why runners are a complement to, not a replacement for, radio and satellite links: incident commanders route urgent, short status updates (casualty counts, "send more ambulances") over radio, and reserve runners for detailed information, formal written orders, or the specific isolated pockets that radio simply cannot reach.
Coordinated Common Operating Picture Restored
The goal of every fallback tier — radio, satellite, mesh, runner — is not communication for its own sake, but a shared, current picture of the disaster: how many patients are where, which hospitals have capacity, which roads are open. In Incident Command System terms this is the Common Operating Picture (COP), and a degraded mixed network can rebuild it even when no single channel is fast or complete on its own.
- Standard: ICS 201 briefing form (incident action plan / COP baseline)
- 2–10 min: Typical COP refresh (recovered) (mixed-channel network)
- Nationwide: NIMS adoption (US federal/state/local standard)
- 2000s–: Post-9/11 interoperability push (FirstNet, common channels)
The Common Operating Picture concept
A Common Operating Picture is a single, shared understanding of the incident — patient counts, resource status, hazards, and open routes — that every node in the response can see and trust. It is built and maintained through the Incident Command System (ICS), the standardized management structure used across US federal, state, and local response (part of the broader National Incident Management System, NIMS), which defines exactly who reports what to whom, on what form, at what interval — regardless of which physical channel carries the message.
Critically, the COP does not require every node to talk to every other node directly. It requires every node to feed accurate, timely reports into the command structure — over whatever channel is available to that node — and for command to push a synthesized picture back out over whatever channels reach each recipient.
Why a mixed, degraded network still works
The restored network in a real disaster is rarely uniform: some hospitals may have partial cellular service back within hours, some field teams operate only by radio for days, and a few remote posts depend on runners the entire time. What matters for coordination is not that every link is fast — it is that every node has at least one working path, however slow, back to command, and that command has a disciplined process (standard forms, set reporting intervals, message logging) for merging inputs of wildly different speed and format into one coherent picture.
This is precisely why interoperability — the ability of different agencies' radios, forms, and command structures to work together — became the central lesson of 9/11 and a design requirement ever since: a COP is only as good as the weakest agency's ability to feed it and receive from it.
What "restored" actually looks like
Recovery is not a return to the pre-disaster network — it is the point at which the mixed system (surviving cellular, radio nets, satellite links, and runner relays) collectively delivers status updates fast and complete enough for command to make timely triage, transport, and resource-allocation decisions. In practice this often means:
• Portable cell-on-wheels (COW) units and Starlink terminals restoring broadband to key hubs within 24–72 hours • Amateur radio and MARS/RACES nets continuing to carry formal traffic for facilities still offline • Runner relay points gradually decommissioned as radio coverage is patched, but retained as a backup at the edges • A steadily rising Common Operating Picture completeness — the fraction of expected reports actually received on schedule — as the visible measure of coordination recovering, even while the underlying network remains a patchwork rather than a full restoration.
Emergency managers do not wait for full infrastructure restoration to declare coordination "working" — they track COP completeness (the share of expected situation reports actually received) as the real recovery metric, because a degraded mixed network that reliably delivers 90% of reports on time outperforms a "restored" but overloaded cellular network that silently drops calls during the next surge.
This simulation models the coordination of medical assistance during a disaster when communication networks fail, emphasizing effective communication and resource allocation.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install