HomeAI & Machine LearningCliff Walking: SARSA vs Q-Learning

Cliff Walking: SARSA vs Q-Learning

Watch two reinforcement-learning agents — on-policy SARSA and off-policy Q-Learning — learn the classic Cliff Walking task side by side in 3D, and see why Q-Learning finds the risky optimal path along the cliff while SARSA learns a safer detour.

AI & Machine Learning3DAdvanced60 FPS📱 Mobile-adapted⇄ 2D version
ds-topic-33 ↗ Open standalone

Two reinforcement-learning agents race across the same 4×12 "Cliff Walking" gridworld at once — one trained with on-policy SARSA, the other with off-policy Q-Learning — so you can watch the textbook divergence between the two temporal-difference control algorithms happen live. Every step costs a small penalty, falling off the cliff costs a large one and resets the agent to start, and reaching the goal ends the episode. Adjust exploration, learning rate and discount factor and watch each agent's Q-value heat-map reshape itself: Q-Learning converges on the fast, risky path that hugs the cliff edge, while SARSA learns to give the cliff a wider berth because its updates account for its own exploratory mistakes.

⚙ Under the hood

Watch two reinforcement-learning agents — on-policy SARSA and off-policy Q-Learning — learn the classic Cliff Walking gridworld side by side, and see why Q-Learning finds the risky path along the cliff edge while SARSA learns a safer detour.

reinforcement-learningsarsaq-learninggridworldtemporal-differenceon-policy-vs-off-policy

3D · Three.js / WebGL renderer · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)