Chatbot Arena Elo Rating Dynamics
Interactive 3D simulator of the Chatbot Arena / LMSYS pairwise-battle ranking system: watch Elo ratings converge toward hidden true model skill as random head-to-head battles are resolved, and tune K-factor, battle rate and skill spread live.
Modern LLM leaderboards like Chatbot Arena (LMSYS), rather than grading models against a fixed answer key the way MMLU or HumanEval do, rank them purely from crowd-sourced pairwise battles — "which response was better?" — and feed the outcomes into an Elo/Bradley-Terry rating update, the same mathematics chess ratings use. This simulator gives every model in the arena a hidden true skill you can never see directly, runs random head-to-head battles between them with a real logistic win probability, and updates each model's visible Elo rating live in 3D as the bars rise and fall. Tune the K-factor to see the stability/responsiveness trade-off, widen or narrow the hidden skill spread to see how much harder close matchups are to rank, and watch the Spearman rank-correlation readout climb toward 1.0 as enough battles accumulate to recover the true ordering — exactly the statistical process that makes a crowd-voted leaderboard converge to something trustworthy.
Interactive 3D simulator of the Chatbot Arena / LMSYS pairwise-battle ranking system: random head-to-head battles between models with hidden true skill update visible Elo ratings live, showing how a crowd-voted leaderboard converges toward the real ranking.
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install