When a backend blips, every mobile client that was connected notices at roughly the same instant and tries to reconnect at roughly the same instant — a thundering herd. If they all retry on a fixed schedule, the server sees a periodic hammer of traffic and can never drain the queue. Each client's retry delay after attempt n follows exponential backoff, capped at a maximum:
cap(n) = min(maxDelay, base · 2ⁿ)
The three strategies you can switch between decide how the delay is drawn from that cap:
None: delay = cap(n)
Full Jitter: delay = random(0, cap(n))
Decorrelated: delay = min(maxDelay, random(base, prevDelay · 3))
Each simulated server tick, every client whose timer has fired sends one request. The server can only accept capacity requests per tick — the rest fail and re-arm with a longer, randomized delay. None keeps every client's retries in lock-step, so the herd re-synchronizes every cycle and the server load oscillates between empty and overloaded. Full Jitter spreads attempts uniformly across the whole window, draining the herd fastest with the least peak load. Decorrelated Jitter spreads them out even further as attempts accumulate, trading a slightly slower initial drain for even less correlation between clients.
This is the exact mechanism behind retry logic in real mobile SDKs (AWS, Firebase, gRPC) reconnecting after a network handover, an API rate-limit response, or a backend restart — badly-designed retries are one of the most common causes of self-inflicted backend outages.
This 2D version renders the same client/server model top-down — clients ring a central server node, and colored pulses show requests in flight.