Active instance Booting (not yet serving) Draining (finishing requests)
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

Auto-Scaling Group Simulator

This simulation makes cloud auto-scaling policy something you can watch and tune rather than read about. A fleet of up to sixteen instances sits on a 3D rack; incoming request traffic drifts and spikes, utilization is computed against the fleet's total capacity, and a scaling policy launches new instances when utilization stays above a threshold or drains and retires instances when demand falls. Every action is bound by realistic mechanics — new instances take real boot time before they can serve traffic, retiring instances drain their in-flight connections first, and a cooldown timer prevents the fleet from flapping — so you can directly see the trade-off between over-provisioning (higher cost, near-zero queue) and under-provisioning (lower cost, requests queuing during a spike). Live readouts track instance count, utilization, queue depth and estimated hourly cost.