HomeMachine Learning & Neural NetworksMixture of Experts: Sparse Gated Routing (2D)

Mixture of Experts: Sparse Gated Routing (2D)

Interactive 2D mixture-of-experts router: tokens are noisily scored and routed to their top-k experts on a rotatable radial diagram, alongside a live score panel, with an auxiliary-loss-free load-balancing bias you can toggle and expert-utilization readouts.

Machine Learning & Neural Networks2DAdvanced60 FPS📱 Mobile-adapted⇄ 3D version
2d-ds-topic-22 ↗ Open standalone

Every glowing dot leaving the gate at the center of the radial diagram is a token being routed to the small set of "expert" sub-networks the noisy top-k gate selected for it — exactly the sparse-activation mechanism that lets modern large language models pack far more parameters into a model than they ever touch per token. Drag the diagram to spin the expert ring, watch each token's raw vs. noisy score race in the panel below, tune top-k and the gating noise to see routing sharpen or blur, and flip on the auxiliary-loss-free load-balancing bias — the same trick used in DeepSeek-V3 — to watch overloaded experts get throttled back until utilization evens out across all six.

⚙ Under the hood

Watch tokens leave a gating network and travel across a rotatable 2D radial diagram to the top-k expert sub-networks a noisy top-k router selects for them, with a live score panel and an auxiliary-loss-free load-balancing bias you can toggle to flatten expert utilization in real time.

mixture-of-expertsneural-networksdeep-learninggating-networkload-balancingtransformers

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)