Robot Goal Obstacle Safety envelope
Cumulative crashesnaivesafe
Reward / episode (avg)naivesafe
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

Safe RL: Constrained Exploration vs. Naive Exploration

A side-by-side reinforcement-learning training race. Both robots learn the same grid-navigation task from scratch by trial and error. The naive agent explores unconstrained and racks up crashes from falling off the table edge and colliding with the obstacle before its policy improves. The safe agent maintains a learned safety envelope around hazards and reroutes any risky exploratory or greedy action into a safe fallback move, learning just as effectively with zero crashes.