Rating axis — dot position = live EloSpearman ρ convergence over battles
⚠ Couldn't draw the simulatorSomething went wrong initializing the canvas. Reload the page.

Chatbot Arena Elo Rating Dynamics (2D)

Modern LLM leaderboards like Chatbot Arena (LMSYS), rather than grading models against a fixed answer key, rank them purely from crowd-sourced pairwise battles — "which response was better?" — and feed the outcomes into an Elo/Bradley-Terry rating update, the same mathematics chess ratings use. This 2D simulator gives every model in the arena a hidden true skill you can never see directly, runs random head-to-head battles between them with a real logistic win probability, and slides each model's visible Elo rating live along a horizontal axis as battles resolve. A second panel plots the Spearman rank-correlation between the board's current order and the hidden true-skill order as a running curve across the whole battle history, so you watch the accuracy of a crowd-voted leaderboard climb toward 1.0 in real time — exactly the statistical process that makes Chatbot Arena's ranking trustworthy after enough community votes.