Bar height = current Elo rating Pulses = model just battled
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

Chatbot Arena Elo Rating Dynamics

Modern LLM leaderboards like Chatbot Arena (LMSYS), rather than grading models against a fixed answer key the way MMLU or HumanEval do, rank them purely from crowd-sourced pairwise battles — "which response was better?" — and feed the outcomes into an Elo/Bradley-Terry rating update, the same mathematics chess ratings use. This simulator gives every model in the arena a hidden true skill you can never see directly, runs random head-to-head battles between them with a real logistic win probability, and updates each model's visible Elo rating live in 3D as the bars rise and fall. Tune the K-factor to see the stability/responsiveness trade-off, widen or narrow the hidden skill spread to see how much harder close matchups are to rank, and watch the Spearman rank-correlation readout climb toward 1.0 as enough battles accumulate to recover the true ordering — exactly the statistical process that makes a crowd-voted leaderboard converge to something trustworthy.