Agent Benchmark Significance Lab (2D)
Run two tool-using LLM agents through a seeded task battery and see whether one really is better: a permutation test on the success-rate gap plus a bootstrap confidence interval, rendered as a 2D pass/fail grid, a draggable null-distribution histogram and a bootstrap CI strip.
Two tool-using LLM agents each run a seeded battery of pass/fail tasks, and the question this simulator answers is the one every real evaluation harness has to answer: is the difference in their success rates real, or could it just as easily be noise? This 2D dashboard renders both agents' task outcomes as pass/fail grids, runs a permutation test that reshuffles which trials "belong" to which agent thousands of times to build an empirical null distribution (drag to pan, scroll to zoom), and a bootstrap resample to produce a 95% confidence interval on the success-rate gap — the same resampling methodology used to validate benchmark claims before trusting them.
Run two tool-using LLM agents through a seeded task battery and see whether one really is better: a permutation test on the success-rate gap plus a bootstrap confidence interval, rendered as a 2D pass/fail grid, a draggable/zoomable null-distribution histogram and a bootstrap CI strip.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install