🧠 Evaluating Reasoning in LLMs: Benchmark & Chain-of-Thought Verification Lab
Watch an LLM benchmark grid get scored live: tune task difficulty, chain-of-thought length, tool-assisted verification and self-consistency sampling, and see accuracy, compute cost and seed variance respond in real time.
AI & Machine Learning3DModerate60 FPS
⚙ Under the hood
Watch an LLM benchmark grid get scored live: tune task difficulty, chain-of-thought length, tool-assisted verification and self-consistency sampling, and see accuracy, compute cost and seed variance respond in real time.
Three.jsLLM reasoningchain-of-thoughtbenchmark evaluationInstancedMesh
3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install