Adaptive Labyrinth Simulator (2D)
Interactive 2D Q-learning maze: watch a reinforcement-learning agent improve its route over repeated episodes, visualized as a live state-value heatmap and greedy policy arrows.
A tabular Q-learning agent explores a top-down grid maze and, episode after episode, adapts its route using the Bellman update Q(s,a) += α·[r + γ·max Q(s′,·) − Q(s,a)]. A live heatmap shows the learned state values and arrows trace the greedy policy as it sharpens from random wandering into a direct path — and every ~40 episodes the maze itself reshuffles a few walls, so the agent must keep adapting rather than memorize one route.
Interactive 2D Q-learning maze: watch a reinforcement-learning agent improve its route over repeated episodes, visualized as a live state-value heatmap and greedy policy arrows.
2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install