Home▸AI & Machine Learning▸Evaluating Reasoning in LLMs (2D)

Evaluating Reasoning in LLMs (2D)

2D canvas version: tune task difficulty, chain-of-thought length, tool-assisted verification and self-consistency sampling, and watch a benchmark grid plus a live reasoning-chain walk score in real time.

AI & Machine Learning2DModerate60 FPS📱 Mobile-adapted⇄ 3D version
2d-evaluating-reasoning-in-llms-benchmark-chain-of-thought ↗ Open standalone

This 2D canvas companion drives the same chain-of-thought scoring model as the 3D version: per-step correctness p = 1 − difficulty, chain probability p^steps, an optional tool-assisted verifier that catches ~55% of chain errors, and a self-consistency majority vote across k samples. A colored grid of benchmark questions resamples every few seconds, and a chain-of-thought walk above it plays out one example reasoning chain step by step.

⚙ Under the hood

2D canvas version: tune task difficulty, chain-of-thought length, tool-assisted verification and self-consistency sampling, and watch a benchmark grid plus a live reasoning-chain walk score in real time.

LLM reasoningchain-of-thoughtbenchmark evaluationself-consistency2D

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)