Task passed Task failed Null-distribution bar Observed gap / CI bounds
drag histogram to pan · scroll to zoom

Agent Benchmark Significance Lab (2D)

Two tool-using LLM agents each run a seeded battery of pass/fail tasks, and the question this simulator answers is the one every real evaluation harness has to answer: is the difference in their success rates real, or could it just as easily be noise? This 2D dashboard renders both agents' task outcomes as pass/fail grids, runs a permutation test that reshuffles which trials "belong" to which agent thousands of times to build an empirical null distribution (drag to pan, scroll to zoom), and a bootstrap resample to produce a 95% confidence interval on the success-rate gap — the same resampling methodology used to validate benchmark claims before trusting them.