Requests arrive as a Poisson process with rate λ: in each fixed timestep dt the number of new arrivals is drawn from Poisson(λ·dt) (Knuth's algorithm). Every server instance behaves as one channel of an M/M/c queue — while busy it completes its current request with probability 1−e−μdt per step, the exact discretization of an exponential service time with rate μ. Idle servers pull the next waiting request FIFO the instant they free up.
arrivals(dt) ~ Poisson(λ·dt)
P(complete|dt) = 1 − e^(−μ·dt) (per busy server)
ρ (utilization)= busy_servers / total_servers
scale up: ρ > up_threshold → +1 server (cooldown 3s)
scale down: ρ < down_threshold → −1 idle server (cooldown 3s)
Little's Law: L = λ·W (avg-in-system = throughput × avg time-in-system)
The auto-scaler is the same reactive rule real cloud platforms use (e.g. AWS/GCP target-utilization scaling): it samples average utilization over a rolling window and adds or removes one instance at a time, with a cooldown so it doesn't flap on noise. A capped queue (300 requests) models admission control — once full, new arrivals are dropped (503) instead of queueing forever. The Little's-Law readout compares the directly measured average queue+in-service population L against λ·W computed from measured throughput and response time — in a stable system the two track each other closely, which is the classic sanity check for any queueing simulation.
- Server grid — each cell is one instance; filled = busy, outline = idle; the pool grows/shrinks live as the autoscaler acts.
- Queue strip — dots waiting for a free server, capped at 300 (overflow drops requests).
- Charts — queue length & server count over time, and average response time / utilization over time.