Modern LLM leaderboards like Chatbot Arena (LMSYS), rather than grading models against a fixed answer key the way MMLU or HumanEval do, rank them purely from crowd-sourced pairwise battles — "which response was better?" — and feed the outcomes into an Elo/Bradley-Terry rating update, the same mathematics chess ratings use. This simulator gives every model in the arena a hidden true skill you can never see directly, runs random head-to-head battles between them with a real logistic win probability, and updates each model's visible Elo rating live in 3D as the bars rise and fall. Tune the K-factor to see the stability/responsiveness trade-off, widen or narrow the hidden skill spread to see how much harder close matchups are to rank, and watch the Spearman rank-correlation readout climb toward 1.0 as enough battles accumulate to recover the true ordering — exactly the statistical process that makes a crowd-voted leaderboard converge to something trustworthy.