Agent
Wall
Goal
High value
A tabular Q-learning agent explores a top-down grid maze and, episode after episode, adapts its route using the Bellman update Q(s,a) += α·[r + γ·max Q(s′,·) − Q(s,a)]. A live heatmap shows the learned state values and arrows trace the greedy policy as it sharpens from random wandering into a direct path — and every ~40 episodes the maze itself reshuffles a few walls, so the agent must keep adapting rather than memorize one route.