HomeAI & Machine LearningInference Autoscaling: Queueing Theory for ML Model Serving

Inference Autoscaling: Queueing Theory for ML Model Serving

Interactive M/M/c queueing simulator for ML model serving: watch Poisson-arriving inference requests queue for replicas, tune arrival rate, service rate and target utilization, and see a Kubernetes-style autoscaler add or remove replicas live.

AI & Machine Learning3DModerate60 FPS📱 Mobile-adapted⇄ 2D version
ds-topic-62 ↗ Open standalone

Every production ML model behind an API sits inside an M/M/c queue whether its operators think about it that way or not: requests arrive at some rate λ, a pool of replicas each serve them at rate μ, and the gap between those two numbers decides whether latency stays flat or explodes. This simulator renders a live inference fleet in 3D — request particles spawn, queue for a free replica, get processed and either complete or get dropped when the wait buffer overflows — while a Kubernetes-style autoscaler watches the fleet's utilization ρ = λ/(c·μ) and adds or removes replicas to hold it near your chosen target. Tune the arrival rate, per-replica service rate and target utilization and watch p95 latency, queue depth, dropped requests and estimated hourly cost respond in real time, exactly the trade-off behind every HPA policy and SageMaker/Vertex AI autoscaling config in production.

⚙ Under the hood

An interactive M/M/c queueing simulator for ML model serving — tune arrival rate, per-replica service rate and autoscaler target utilization while a Kubernetes-style autoscaler adds or removes replicas live, and watch p95 latency, queue depth, dropped requests and cost respond.

mlopsautoscalingqueueing-theorymodel-servingdevopskubernetes

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)