This simulation demonstrates the core mechanics of serverless computing — applications that run without provisioning or managing servers directly. Incoming events stream into a queue at a controllable rate; an auto-scaler continuously computes how many ephemeral function instances are needed to keep up, spinning up new ones (paying a real cold-start latency penalty) or tearing down instances that have sat idle past a timeout, exactly as platforms like AWS Lambda, Google Cloud Run and Knative do in production. Live readouts track queue depth, active vs. cold-starting capacity, average request latency and accrued compute cost as you tune arrival rate, per-instance concurrency, cold-start latency and the idle scale-in timeout — or trigger a traffic spike to watch the whole event-driven, auto-scaling system react in real time.