drag to rotate

Mixture of Experts: Sparse Gated Routing (2D)

Every glowing dot leaving the gate at the center of the radial diagram is a token being routed to the small set of "expert" sub-networks the noisy top-k gate selected for it — exactly the sparse-activation mechanism that lets modern large language models pack far more parameters into a model than they ever touch per token. Drag the diagram to spin the expert ring, watch each token's raw vs. noisy score race in the panel below, tune top-k and the gating noise to see routing sharpen or blur, and flip on the auxiliary-loss-free load-balancing bias — the same trick used in DeepSeek-V3 — to watch overloaded experts get throttled back until utilization evens out across all six.