This simulation models the exact trade-off a cloud cost engineer faces every day: how aggressively should a fleet autoscale? A synthetic workload — diurnal, steady, or spiky — drives demand for compute capacity, and a target-utilization autoscaling policy decides how many virtual-machine instances to run, gated by a realistic scale-action cooldown. Every instance is rendered as a lit 3D server block whose colour reflects its load, while live readouts track active instance count, fleet utilization, cost per hour, cumulative spend, and the share of simulated time spent in SLA violation (demand exceeding provisioned capacity). Push the target utilization up and watch cost fall but SLA breaches climb; lower it and the fleet over-provisions for safety — the same lever real infrastructure teams pull to balance performance against expense.