Model Serving Architectures for Production ML (2D)
2D inference-cluster lab: tune request rate, batch size, replica count and autoscaling to watch queueing latency, throughput and queue depth respond in real time.
This 2D companion runs the same queueing math as the 3D cluster on a flat canvas: particles stream from a client node through a load balancer to a row of replica boxes, each replica reddening as its load climbs. A latency sparkline along the top makes the classic utilization "hockey stick" easy to read at a glance.
⚙ Under the hood
2D inference-cluster lab: tune request rate, batch size, replica count and autoscaling to watch queueing latency, throughput and queue depth respond in real time.
model servinginferenceautoscalingdynamic batchingqueueing theorylatency
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install