Tower A Tower B (verbose) Correct verdict Biased / wrong verdict
drag to scrub history

LLM-as-a-Judge: Position Bias Lab (2D)

Using a large language model to score or compare other models' outputs — "LLM-as-a-judge" — has become the default way to evaluate chain-of-thought reasoning, RLHF preference pairs and chatbot leaderboards, but the judge itself carries systematic biases. This 2D simulator runs a stream of pairwise trials between response A and a deliberately longer response B, mixing a true quality signal with position bias (favoring whichever response is shown first) and verbosity bias (favoring the longer one), racing particles across a flat arena toward the winning tower and charting accuracy / position-error live in a scrollable trend panel. Tune the bias strengths, watch the two real mitigations respond live: self-consistency ensembling, which shrinks noise but leaves systematic bias untouched, versus order counterbalancing, which cancels position bias from the decision by construction.