GPU node Reduced gradient chunk In-flight chunk (ring link)
drag to rotate
Compute vs comm — last 40 steps

Ring All-Reduce 2D: Bandwidth vs Latency

The 2D companion to the 3D ring all-reduce simulator renders the same GPU cluster gradient-sync mechanism as a flat, draggable ring diagram paired with a live compute-vs-communication timeline strip chart. Tune cluster size, model parameter count, interconnect bandwidth, per-GPU compute throughput and — new in this build — per-hop network latency, then run a training step to watch scatter-reduce and all-gather chunks circulate the ring while the timeline logs whether the cluster is compute-bound, latency-bound or bandwidth-bound.