Each running instance can serve a fixed capacity of 40 requests/second. With N serving instances (active + still-draining) the fleet's utilization is:
u(t) = RPS(t) / (N(t) × 40)
queue' = max(0, RPS − N×40) // grows when demand exceeds capacity
queue = max(0, queue − max(0, N×40 − RPS)·dt) // drains when there's spare
This mirrors a real auto-scaling group (AWS ASG, GCP MIG, Kubernetes HPA): a scaling policy watches utilization, not the raw request count. If u > threshold holds for a short sustained window, the group launches a new instance — which needs several seconds of real boot latency (bootstrapping, health checks) before it starts serving, so capacity lags demand and the queue absorbs the gap. If u falls under half the threshold for the same sustained window, the newest instance is put into connection draining — it keeps serving in-flight traffic for a few seconds before terminating, avoiding dropped requests.
After every scale-out or scale-in decision, a cooldown timer locks out further scaling actions. Too short a cooldown causes flapping (launching and killing instances back-to-back); too long a cooldown leaves the fleet unable to react to the next spike — the classic operational trade-off this simulator makes visible.
- Incoming traffic — the base request rate driving demand (with a slow sine drift plus noise, like real diurnal + jitter traffic).
- Scale-out threshold — lower it to over-provision (fleet stays large, cost rises, queue stays near zero); raise it to under-provision (fewer instances, lower cost, but the queue spikes and requests wait longer whenever traffic jumps).
- Cooldown period — how long the fleet must wait after any scaling action before it can scale again.
- Traffic spike — injects a 3× burst for 10 seconds so you can watch boot latency create a temporary queue backlog even though the policy reacted correctly.