Each active replica can process capacity = base × (1 + 0.1×(batch−1)) requests/second — bigger batches raise per-replica throughput. Utilization is rate / (capacity × replicas). As utilization approaches 1, an M/M/1-style queueing factor 1 / (1 − util) makes latency and queue depth blow up — the classic "hockey stick." With autoscaling on, the replica count chases a target utilization of 60%, capped at "Max replicas."